Papers by Jinsong Ni
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval (2025.acl-long)
Copied to clipboard
| Challenge: | Document retrieval in real-world scenarios faces significant challenges due to diverse document formats and modalities. |
| Approach: | They propose a visual-textual embedding framework that integrates textual and visual features for robust document representation. |
| Outcome: | The proposed visual-textual embedding framework surpasses existing methods while preserving semantic fidelity. |