Challenge: Multimodal embedding models encode multimedia inputs into latent vector representations.
Approach: They propose to synthesize multimodal multilingual data using a multimodal large language model . they identify three criteria for high-quality synthetic multimodal data .
Outcome: The proposed model outperforms existing models on the MMEB Benchmark and the XTD benchmark.

Similar Papers

Improving Text Embeddings with Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for obtaining text embeddings require complex training pipelines . authors leverage proprietary LLMs to generate diverse synthetic data for text embeds based on 93 languages .
Approach: They propose a method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps.
Outcome: The proposed method achieves strong performance on competitive text embedding benchmarks without using any labeled data.
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct (2025.findings-acl)

Copied to clipboard

Challenge: a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling .
Approach: They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution.
Outcome: The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data.
Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text (2026.findings-eacl)

Copied to clipboard

Challenge: Existing multimodal models rely on machine translation, but performance drops for other languages due to limited multilingual multimodal resources.
Approach: They propose a lightweight alignment method that learns only a few linear layers using English text alone to map multilingual text embeddings into multimodal space.
Outcome: M2M achieves strong zero-shot transfer on XTD Text-to-Image retrieval in English and spanish . it learns only a few linear layers to map multilingual text embeddings into multimodal space .
Train a Unified Multimodal Data Quality Classifier with Synthetic Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal Large Language Models are pre-trained on image-text caption data and interleaved document data.
Approach: They propose to train an efficient MLLM as a Unified Mulitmodal Data Quality Classifier to filter image-text caption and interleaved data.
Outcome: The proposed method enables efficient creation of sample-score pairs for caption and interleaved data to train UniFilter.
Understanding the Influence of Synthetic Data for Text Embedders (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in general purpose text embedders have been driven by training on synthetic training data.
Approach: They propose to use GPT-4 to produce high quality synthetic data that expands existing training datasets for embeddings to new tasks.
Outcome: The proposed dataset is high quality and leads to consistent improvements in performance.
MM-LLMs: Recent Advances in MultiModal Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: MultiModal Large Language Models (MM-LLMs) have undergone significant advances in the past year . traditional MM models incur substantial computational costs, especially when trained from scratch .
Approach: They propose a taxonomy encompassing 126 MM-LLMs and summarize key training recipes to enhance their potency.
Outcome: The proposed models preserve the reasoning and decision-making capabilities of LLMs and empower diverse range of MM tasks.
Little Giants: Synthesizing High-Quality Embedding Data at Scale (2025.naacl-long)

Copied to clipboard

Challenge: Synthetic data generation is an increasingly popular way of training models without the need for large, manually labeled datasets.
Approach: They propose a framework that aligns open-source small models to efficiently generate large-scale embedding data.
Outcome: The proposed framework outperforms state-of-the-art embedding models by using only 1/10 of the GPT API calls.
Beyond Triplet: Leveraging the Most Data for Multimodal Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision.
Approach: They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset.
Outcome: The proposed method can significantly improve translation performance with more non-triple data.
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in embedding resources have led to a lack of representation of the Dutch language in multilingual resources.
Approach: They introduce Massive Text Embedding Benchmark for Dutch (MTEB-NL) which includes existing Dutch datasets and newly created ones, covering a wide range of tasks.
Outcome: The proposed models demonstrate strong performance across multiple tasks.
Retrieving Multimodal Information for Augmented Generation: A Survey (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly using multimodality to augment their generation ability, but there is no unified perception of at which stage and how to incorporate different modalities.
Approach: They propose to use multimodality to augment Large Language Models (LLMs) this will provide scholars with a deeper understanding of the methods' applications and encourage them to adapt existing techniques to the fast-growing field of LLMs.
Outcome: The proposed methods improve factuality, reasoning, interpretability, and robustness of the generated content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations