Challenge: Existing models implicitly assume that documents in different languages are highly comparable, a false assumption.
Approach: They propose a multilingual topic model that learns weighted topic links and connects cross-lingual topics only when the dominant words defining them are similar.
Outcome: The proposed model outperforms existing models in low-resource language tasks and outperformed LDA and previous models in classification tasks using documents’ topic posteriors as features.

Similar Papers

Learning Multilingual Topics from Incomparable Corpora (C18-1)

Copied to clipboard

Challenge: Existing models require parallel or comparable training, which limits their ability to generalize.
Approach: They propose a method that demystifies the knowledge transfer mechanism behind multilingual topic models by defining an alternative but equivalent formulation.
Outcome: The proposed model learns coherent multilingual topics from partially and fully incomparable corpora with limited amounts of dictionary resources.
LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing cross-lingual topic models depend on sparse bilingual resources and often yield incoherent or weakly aligned topics.
Approach: They propose a framework that integrates LLM-guided topic refinement with self-consistency uncertainty quantification to enable black-box, stable, and scalable enhancement of cross-lingual topic models.
Outcome: Experiments on multilingual corpora show that the proposed framework achieves superior topic coherence and alignment while reducing reliance on bilingual dictionaries and expensive LLM calls.
Lessons from the Bible on Modern Topics: Low-Resource Multilingual Topic Model Evaluation (N18-1)

Copied to clipboard

Challenge: Existing metrics to evaluate multilingual topic quality are inadequate for multilingual document analysis.
Approach: They propose a new intrinsic evaluation metric for multilingual topic models that correlates well with human judgments of multilingual coherence and performance in downstream applications.
Outcome: The proposed model improves the performance of multilingual topic models in low-resource languages and with human judgments of multilinguistic topic coherence.
Cross-lingual Contextualized Topic Models with Zero-shot Learning (2021.eacl-main)

Copied to clipboard

Challenge: Existing topic models are language-specific and cannot be transferred in a transferable manner.
Approach: They propose a zero-shot cross-lingual topic model that learns topics on one language and predicts them for unseen documents in different languages.
Outcome: The proposed model learns topics on one language and predicts them for unseen documents in different languages.
XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignments (2025.findings-emnlp)

Copied to clipboard

Challenge: XTRA aims to uncover shared semantic themes across languages . previous methods have achieved improvements in topic diversity but struggle to ensure high topic coherence and consistent alignment across languages.
Approach: a new framework unifies Bag-of-Words modeling with multilingual embeddings is proposed to address this problem . XTRA introduces two core components: (1) representation alignment and (2) topic alignment to enforce cross-lingual consistency.
Outcome: XTRA outperforms baselines in topic coherence, diversity, and alignment quality on multilingual corpora.
Multilingual and Multimodal Topic Modelling with Pretrained Embeddings (2022.coling-1)

Copied to clipboard

Challenge: a novel neural topic model for comparable data maps texts from multiple languages and images into a shared topic space.
Approach: They propose a novel multimodal multilingual neural topic model that maps texts from multiple languages and images into a shared topic space.
Outcome: The proposed model outperforms a zero-shot topic model in predicting topic distributions for comparable multilingual data and performs as well on unaligned embeddings as it does on aligned embeds.
ProtoXTM: Cross-Lingual Topic Modeling with Document-Level Prototype-based Contrastive Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study has demonstrated that cross-lingual topic modeling can extract aligned and semantically coherent topics from bilingual corpora.
Approach: They propose a document-level prototype-based contrastive learning paradigm for cross-lingual topic modeling .
Outcome: The proposed approach achieves state-of-the-art performance on cross-lingual and mono-lingual benchmarks.
Multilingual and cross-lingual document classification: A meta-learning approach (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods to document classification in low-resource languages are under-resourced . 6% of the world's languages are spoken, and many have inadequate resources .
Approach: They propose a meta-learning approach to document classification in low-resource languages . they propose 'nuclear-shot' cross-lingual adaptation to previously unseen languages based on limited data .
Outcome: The proposed method performs on-par on some languages while under-resourced in others.
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones .
Approach: They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models.
Outcome: The proposed model consistently outperforms state-of-the-art models across tasks and settings.
Neural Topic Modeling with Large Language Models in the Loop (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, but their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency.
Approach: They propose a novel LLM-in-the-loop framework that integrates Large Language Models with Neural Topic Models (NTMs) global topics and document representations are learned through the NTM, while an LLM refines these topics using an Optimal Transport (OT)-based alignment objective.
Outcome: The proposed framework improves topic interpretability while preserving the efficiency of existing NTMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations