Papers by Truong Nguyen

11 papers
EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignments (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for knowledge distillation focus on direct output alignment, neglecting this crucial structural information.
Approach: They propose a framework for knowledge distillation that maps tokens one-to-one and aligns attention matrix patterns using Centered Kernel Alignment.
Outcome: The proposed framework significantly outperforms existing CTKD baselines.
MIPIC: Matryoshka Representation Learning via Self-Distilled Intra-Relational and Progressive Information Chaining (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to train dense representations require explicit coordination of how information is arranged across embedding dimensionality and model depth.
Approach: They propose a framework that trains Matryoshka representations using self-distilled intra-relational alignment and Progressive information chaining.
Outcome: The proposed framework produces coherent and compact Matryoshka representations with significant performance advantages under low-dimensional models.
FAID: Fine-grained AI-generated Text Detection using Multi-task Auxiliary and Multi-level Contrastive Learning (2026.eacl-long)

Copied to clipboard

Challenge: Existing binary detection frameworks for human-written, LLM-generated and human-LLM collaborative texts are challenging . a recent study focused on binary detection, i.e., human vs. LLM, or on fine-grained detection limited to English.
Approach: They propose a fine-grained detection framework to classify text into three categories . they use multilingual datasets and a multi-domain, multi-generator dataset .
Outcome: The proposed framework outperforms baselines on unseen domains and new LLMs.
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology (2025.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art multimodal language models (MLMs) show promise for supporting SLPs, but their use remains underexplored due to a limited understanding of their performance in high-stakes clinical settings.
Approach: They propose a taxonomy of real-world use cases of multimodal language models in speech-language pathologies to address this gap.
Outcome: The proposed model outperforms 15 state-of-the-art models in speech-language pathologies across five use cases and achieves improvements of over 30% on domain-specific data.
Crossing Linguistic Horizons: Finetuning and Comprehensive Evaluation of Vietnamese Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing open-source LLMs exhibit limited effectiveness in processing Vietnamese . lack of systematic benchmark datasets and metrics tailored for Vietnamese LLM evaluation exacerbates these issues.
Approach: They propose to fine tune LLMs specifically for Vietnamese and develop a framework for evaluation . they find that larger models introduce more biases and uncalibrated outputs .
Outcome: The proposed framework finetunes LLMs specifically for Vietnamese and provides a framework for evaluation .
CovRelex-SE: Adding Semantic Information for Relation Search via Sequence Embedding (2023.eacl-demo)

Copied to clipboard

Challenge: COVID-19 has affected all aspects of human life, causing problems related to acronyms, synonyms, and rare keywords.
Approach: They propose a hybrid relation retrieval system based on embeddings to provide high-quality search results.
Outcome: The proposed system can be accessed through the following URL: http://www.jaist.ac.jp/is/labs/nguyen-lab/systems/covrelex-se/.
Automated Generation of Accurate & Fluent Medical X-ray Reports (2021.emnlp-main)

Copied to clipboard

Challenge: Existing medical report generation efforts focus on producing human-readable reports, yet the generated text may not be well aligned to the clinical facts.
Approach: They propose to automate the generation of medical reports from chest X-ray image inputs . medical reports are the primary medium, which physicians communicate findings from scans - authors say .
Outcome: The proposed method achieves fluency and clinical accuracy on common metrics.
COVID-19 Named Entity Recognition for Vietnamese (2021.naacl-main)

Copied to clipboard

Challenge: a new dataset is being developed to help fight the COVID-19 pandemic . the dataset is annotated for the named entity recognition task with newly-defined entity types .
Approach: They present the first manually-annotated COVID-19 domain-specific dataset for Vietnamese . their dataset is annotated for the named entity recognition task with newly-defined entity types .
Outcome: The proposed dataset is the first manually-annotated COVID-19 domain-specific dataset for Vietnamese.
StructSP: Efficient Fine-tuning of Task-Oriented Dialog System by Using Structure-aware Boosting and Grammar Constraints (2023.findings-acl)

Copied to clipboard

Challenge: Existing models that learn hierarchical structure information representations do not perform well on task-oriented dialog systems.
Approach: They propose a hierarchical structure information representation model that reinforces the semantic awareness of a pre-trained language model by a two-step fine-tuning mechanism.
Outcome: The proposed model is better than existing models at learning the contextual representations of utterances embedded within its hierarchical semantic structure and improves system performance.
ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining (2022.lrec-1)

Copied to clipboard

Challenge: Recent large-scale language models show remarkable achievements in key NLP tasks such as Question Answering and Text Summarization.
Approach: They propose a domain-specific pre-trained Vietnamese language model that outperforms the general domain language models.
Outcome: The proposed model outperforms the general domain language models in Vietnamese datasets while outperforming the general-domain language models.
HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies suggest that fixing the routers can achieve competitive performance by alleviating the collapsing problem, where all experts eventually learn similar representations.
Approach: They propose a method that dynamically generates router parameters through a fixed hypernetwork and trainable embeddings to achieve a balance between training the routers and freezing them to learn an improved routing policy.
Outcome: Experiments on a wide range of tasks show that the proposed method performs better than existing methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations