Challenge: a novel neural topic model for comparable data maps texts from multiple languages and images into a shared topic space.
Approach: They propose a novel multimodal multilingual neural topic model that maps texts from multiple languages and images into a shared topic space.
Outcome: The proposed model outperforms a zero-shot topic model in predicting topic distributions for comparable multilingual data and performs as well on unaligned embeddings as it does on aligned embeds.

Similar Papers

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text (2026.findings-eacl)

Copied to clipboard

Challenge: Existing multimodal models rely on machine translation, but performance drops for other languages due to limited multilingual multimodal resources.
Approach: They propose a lightweight alignment method that learns only a few linear layers using English text alone to map multilingual text embeddings into multimodal space.
Outcome: M2M achieves strong zero-shot transfer on XTD Text-to-Image retrieval in English and spanish . it learns only a few linear layers to map multilingual text embeddings into multimodal space .
Dynamic Topic Modeling by Clustering Embeddings from Pretrained Language Models: A Research Proposal (2022.aacl-srw)

Copied to clipboard

Challenge: Neural Topic Models (NTMs) are topic models that are created with the help of a pretrained language model.
Approach: They propose to do Neural Topic Modeling by Clustering document Embeddings (NTM-CE) with a pretrained language model to create dynamic topic models.
Outcome: The proposed model can be evaluated theoretically and practically using quantitative measurements of coherence and human evaluation to evaluate the model.
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)

Copied to clipboard

Challenge: a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations .
Approach: They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search .
Outcome: The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX.
A Multilingual Topic Model for Learning Weighted Topic Links Across Corpora with Low Comparability (D19-1)

Copied to clipboard

Challenge: Existing models implicitly assume that documents in different languages are highly comparable, a false assumption.
Approach: They propose a multilingual topic model that learns weighted topic links and connects cross-lingual topics only when the dominant words defining them are similar.
Outcome: The proposed model outperforms existing models in low-resource language tasks and outperformed LDA and previous models in classification tasks using documents’ topic posteriors as features.
m3P: Towards Multimodal Multilingual Translation with Multimodal Prompt (2024.lrec-main)

Copied to clipboard

Challenge: Existing multimodal neural machine translation models focus on bilingual translation, but experimental results show that they outperform the text-only baselines and multilingual multimodal methods by a large margin.
Approach: They propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P) this framework aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation.
Outcome: The proposed framework outperforms previous text-only baselines and multilingual multimodal methods by a large margin.
Multi-source Neural Topic Modeling in Multi-view Embedding Spaces (2021.naacl-main)

Copied to clipboard

Challenge: Recent work has used pre-trained word embeddings to address data sparsity in short-text or small document collections.
Approach: They propose a neural topic modeling framework using multi-view embedding spaces to improve topic quality and deal with polysemy.
Outcome: The proposed framework improves topic quality and deal with polysemy.
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding (2026.acl-long)

Copied to clipboard

Challenge: Recent approaches demonstrate that MLLMs can be adapted into competitive embedding models via large-scale contrastive learning.
Approach: They propose a compressed pre-training phase which serves as a warm-up stage for contrastive learning.
Outcome: The proposed model achieves state-of-the-art among MLLMs of comparable size on the MMEB, realizing optimization in both efficiency and effectiveness.
MCSE: Multimodal Contrastive Learning of Sentence Embeddings (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to learning semantically meaningful sentence embeddings are limited by the complexity of pre-trained models.
Approach: They propose a sentence embedding learning approach that exploits both visual and textual information via a multimodal contrastive objective.
Outcome: The proposed approach improves the state-of-the-art average Spearman’s correlation by 1.7% on a variety of semantic textual similarity tasks.
Cross-lingual Cross-modal Pretraining for Multimodal Retrieval (2021.naacl-main)

Copied to clipboard

Challenge: Recent pretrained vision-language models have achieved impressive performance on cross-modal retrieval tasks in English.
Approach: They propose a new approach to learn cross-lingual cross-modal representations for matching images and captions in multiple languages using an annotated corpus.
Outcome: The proposed model achieves impressive performance on two multimodal multilingual image caption benchmarks: Multi30k with German captions and MSCOCO with Japanese captions.
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal embedding models encode multimedia inputs into latent vector representations.
Approach: They propose to synthesize multimodal multilingual data using a multimodal large language model . they identify three criteria for high-quality synthetic multimodal data .
Outcome: The proposed model outperforms existing models on the MMEB Benchmark and the XTD benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations