Papers by Nguyen Tung
HiCOT: Improving Neural Topic Models via Optimal Transport and Contrastive Learning (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in neural topic models (NTMs) have improved topic quality but still face challenges: weak document-topic alignment, high inference costs due to large pretrained language models, and limited modeling of hierarchical topic structures. |
| Approach: | They propose a framework that integrates hierarchical clustering and contrastive learning to refine document-topic relationships using compact PLM-based embeddings. |
| Outcome: | The proposed framework improves topic coherence, topic performance, representation quality and computational efficiency over existing NTMs. |
SRA: Span Representation Alignment for Large Language Model Distillation (2026.acl-long)
Copied to clipboard
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Trung Le
| Challenge: | Existing approaches to cross-tokenizer knowledge distillation are brittle and sensitive to discrepancies between tokenizers. |
| Approach: | They propose a framework that shifts the unit of alignment from tokens to robust, tokenizer-agnostic spans and employ a geometric regularizer to preserve the structural integrity of the representation space. |
| Outcome: | The proposed framework outperforms state-of-the-art methods in cross-architecture distillation experiments. |
XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignments (2025.findings-emnlp)
Copied to clipboard
| Challenge: | XTRA aims to uncover shared semantic themes across languages . previous methods have achieved improvements in topic diversity but struggle to ensure high topic coherence and consistent alignment across languages. |
| Approach: | a new framework unifies Bag-of-Words modeling with multilingual embeddings is proposed to address this problem . XTRA introduces two core components: (1) representation alignment and (2) topic alignment to enforce cross-lingual consistency. |
| Outcome: | XTRA outperforms baselines in topic coherence, diversity, and alignment quality on multilingual corpora. |
GloCOM: A Short Text Neural Topic Model via Global Clustering Context (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing neural topic models often overlook uncovering hidden topics from short texts due to data sparsity, poor aggregation quality, and difficulty in inferring topic proportions for individual documents. |
| Approach: | They propose a model which constructs global clustering contexts for short texts using text embeddings from pre-trained language models. |
| Outcome: | The proposed model outperforms state-of-the-art models on short texts in topic quality and document representation. |
Sharpness-Aware Minimization for Topic Models with High-Quality Document Representations (2025.naacl-long)
Copied to clipboard
Tung Nguyen, Tue Le, Hoang Tran Vuong, Quang Duc Nguyen, Duc Anh Nguyen, Linh Ngo Van, Sang Dinh, Thien Huu Nguyen
| Challenge: | Recent advances in topic models have significantly enhanced the performance compared to conventional probabilistic approaches. |
| Approach: | They propose to use a topic model optimization method to improve generalization and performance. |
| Outcome: | The proposed method improves topic models on popular benchmark datasets. |
VIMQA: A Vietnamese Dataset for Advanced Reasoning and Explainable Multi-hop Question Answering (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing Vietnamese Question Answering (QA) datasets do not explore the model’s ability to perform advanced reasoning and provide evidence to explain the answer. |
| Approach: | They propose to use Vietnamese as a question-answer dataset with 10,000 Wikipedia-based multi-hop question-and-answ pairs to test model's ability to reason and explain the answer. |
| Outcome: | The proposed dataset is in Vietnamese, a low-resource language. |
Topic Modeling for Short Texts via Optimal Transport-Based Clustering (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to topic modeling are based on probabilistic graphical models or non-negative matrix factorization techniques. |
| Approach: | They propose a method that aligns global clusters with topics to discover topics and learn document representations in topic space. |
| Outcome: | The proposed method outperforms state-of-the-art techniques in short-text topic modeling across commonly used metrics. |
Beyond Coherence: Improving Temporal Consistency and Interpretability in Dynamic Topic Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing topic models capture bag-of-words statistics but lack semantic priors . interpretability remains shallow, relying on noisy top-word lists that obscure thematic clarity. |
| Approach: | They propose a variational framework to capture more faithful temporal trajectories . they propose to use entropy-regularized optimal transport to align entire topic constellations . |
| Outcome: | The proposed framework captures more faithful temporal trajectories and improves interpretability. |
Multi-Surrogate-Objective Optimization for Neural Topic Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Neural topic modeling incorporates multiple loss functions but can be difficult to optimize for disparate magnitudes of these losses. |
| Approach: | They propose a gradient-based multi-objective optimization approach that integrates MOO algorithms into the model without the need for hard-parameter sharing. |
| Outcome: | The proposed approach outperforms direct MOO applications on NTMs. |
VN-MTEB: Vietnamese Massive Text Embedding Benchmark (2026.findings-eacl)
Copied to clipboard
| Challenge: | a lack of large-scale test datasets makes it difficult to evaluate AI models before deploying them in real-world projects. |
| Approach: | They propose a Vietnamese benchmark for embedding models that leverages large language models and embeddable models to translate and filter samples from the Massive Multilingual Text Embedding Benchmark. |
| Outcome: | The proposed benchmark outperforms existing models in Vietnamese and English tasks with 41 datasets. |
NeuroMax: Enhancing Neural Topic Modeling via Maximizing Mutual Information and Group Topic Regularization (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in neural topic models have focused on two main directions: the integration of the inference network with a pre-trained language model and the modeling of the relationship between words and topics in the generative model. |
| Approach: | They propose a framework to maximize mutual information between topics and PLMs . Experimental results indicate that NeuroMax reduces inference time . |
| Outcome: | The proposed framework reduces inference time, generates more coherent topics and topic groups, and produces more representative document embeddings. |
ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to train large language models do not capture how humans learn to think. |
| Approach: | They propose a method to fine-tune large language models for mathematical reasoning by using a text-infilling task that predicts masked equations from a given solution. |
| Outcome: | Experiments on GSM8K, MATH, and GSM-Symbolic show that ClozeMath surpasses baseline Masked Thought in performance and robustness with two test-time scaling decoding algorithms, Beam Search and Chain-of-Thought decoding. |
Building a Personalized Dialogue System with Prompt-Tuning (2022.naacl-srw)
Copied to clipboard
| Challenge: | We build a dialogue system that can respond based on a given character setting (persona) this method is not suitable because the more persona information is added, the longer the input text becomes. |
| Approach: | They propose to use prompt-tuning to build a dialogue system that responds based on a persona . they conduct automatic and manual evaluations on English and Japanese . |
| Outcome: | The proposed method can build a dialogue system with more natural responses with less computational resources than fine-tuning. |
SLM-Bench: A Comprehensive Benchmark of Small Language Models on Environmental Impacts (2025.findings-emnlp)
Copied to clipboard
| Challenge: | SLMs offer computational efficiency and accessibility, yet a systematic evaluation of their performance and environmental impact remains lacking. |
| Approach: | SLM-Bench evaluates 15 SLMs on 9 NLP tasks using 23 datasets . compared accuracy, computational efficiency, and sustainability metrics . |
| Outcome: | SLM-Bench evaluates 15 SLMs on 9 NLP tasks using 23 datasets spanning 14 domains. |