Challenge: Multimodal video retrieval systems are needed for multimodal content retrieval . multimodal video search systems are sub-optimal for multi-modal content representations .
Approach: They propose a model that learns retrieval cues for the textual query from multiple modalities and a shared embedding space with task-specific contrastive loss functions.
Outcome: The proposed model outperforms state-of-the-art methods on the MSR-VTT and YouCook2 datasets and shows significant improvements from baseline.

Similar Papers

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods focus on textual queries that include visual information, but lack the ability to address multimodal queries that encompass both textual and visual information.
Approach: They propose a retrieval framework that achieves modality interaction without fusing textual features during the alignment.
Outcome: The proposed method achieves modality interaction without fusing textual features during the alignment.
Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment (2025.acl-long)

Copied to clipboard

Challenge: Existing studies focus on capturing information in multimodal data that is similar to their paired texts, but often ignores the complementary information contained in multimodule data.
Approach: They propose a multimodal retrieval approach that employs Complementary Information Extraction and Alignment to capture complementary information in multimodal data.
Outcome: The proposed approach achieves significant improvements over divide-and-conquer models and universal dense retrieval models.
Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval (2022.tacl-1)

Copied to clipboard

Challenge: Current approaches to cross-modal retrieval process text and visual input jointly . current approaches are pretrained from scratch and suffer from huge retrieval latency and inefficiency issues .
Approach: They propose a cooperative retrieve-and-rerank framework that turns pretrained text-image multi-modal models into efficient retrieval models.
Outcome: The proposed framework improves retrieval performance over current approaches . it uses twin networks to encode all items of a corpus and a cross-encoder component for a more nuanced ranking .
Cross-Modal Retrieval Augmentation for Multi-Modal Classification (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing.
Approach: They propose a retrieval-augmented multi-modal transformer architecture for embedding images and captions in the same space.
Outcome: The proposed approach improves visual question answering over strong baselines and hot-swapping indices.
Query Generation for Multimodal Documents (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to find relevance for multimodal documents with images are expensive and require a lot of runtime overhead.
Approach: They propose to attach generated queries to doc-uments and index them to narrow down to candidate matches using inverted index.
Outcome: The proposed model improves relevance ranking for multimodal documents with images . the proposed model can achieve the state of the art in the first stage retrieval scenarios .
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data.
Approach: They review training strategies, robustness enhancements, loss functions, and agent-based approaches and outline open challenges and future directions to guide research in this evolving field.
Outcome: The proposed model improves accuracy and accuracy while integrating external dynamic information for improved factual grounding.
Towards Multi-Modal Text-Image Retrieval to improve Human Reading (2021.naacl-srw)

Copied to clipboard

Challenge: In primary school, children's books, as well as in modern language learning apps, multi-modal learning strategies like illustrations of terms and phrases are used to support reading comprehension.
Approach: They propose to use multi-modal transformers to train multi-dimensional models on text-image retrieval to support a user's reading comprehension of arbitrary text.
Outcome: The proposed model performs poorly because of the short and relatively simple textual data that the current models are trained with.
Adaptive Fusion Techniques for Multimodal Data (2021.eacl-main)

Copied to clipboard

Challenge: Effective fusion of data from multiple modalities is challenging due to the heterogeneous nature of multimodal data.
Approach: They propose two adaptive fusion techniques that aim to combine multimodal data effectively.
Outcome: The proposed networks can model context from other modalities better than existing methods.
Revisiting Multimodal Transformers for Tabular Data with Text Fields (2024.findings-acl)

Copied to clipboard

Challenge: Tabular data with text fields can be used in financial risk assessment and diagnosis prediction.
Approach: They propose a tabular/text dual-stream Transformer network with numerical embedding schemes and an overall attention module to estimate whether a prediction is uncertain.
Outcome: The proposed model can estimate whether a prediction is uncertain or not based on two well-informed modality streams .
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively.
Approach: They propose to enable multimodal large language models to memorize and recall images within their parameters.
Outcome: The proposed model performs well even with large-scale image candidate sets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations