Papers by Bach Nguyen
Risk Minimization for Zero-shot Sequence Labeling (2021.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to zero-shot sequence labeling are expensive and hard to obtain for lowresource languages/domains. |
| Approach: | They propose a framework for zero-shot sequence labeling with minimum risk training and a decomposable risk function that models the relations between predicted labels from the source models and the true labels. |
| Outcome: | The proposed framework outperforms state-of-the-art systems on 21 datasets. |
Continual Safety Alignment via Gradient-Based Sample Selection (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models require continuous adaptation to new domains, tasks, and evolving requirements. |
| Approach: | They propose a gradient-based sample selection method that filters high-gradient samples during fine-tuning. |
| Outcome: | The proposed method significantly improves alignment preservation while maintaining competitive task performance on continual domain tasks. |
Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning (2021.acl-long)
Copied to clipboard
| Challenge: | Recent work shows document-level contexts can significantly improve Named Entity Recognition models. |
| Approach: | They propose to find external contexts of a sentence by retrieving and selecting a set of semantically relevant texts through a search engine with the original sentence as the query. |
| Outcome: | The proposed approach can achieve new state-of-the-art performance on 8 NER data sets across 5 domains. |
Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor (2021.acl-long)
Copied to clipboard
Xinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu
| Challenge: | Knowledge distillation is a technique to transfer knowledge between models, typically from a large model (the teacher) to a more fine-grained one (the student). |
| Approach: | They propose a factorized form of the knowledge distillation objective for structured prediction which is tractable for many typical choices of the teacher and student models. |
| Outcome: | The proposed model is able to transfer knowledge between teacher and student models without loss of accuracy under four different scenarios. |
Revisiting Generalization Across Difficulty Levels: It’s Not So Easy (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing research is mixed regarding whether training on easier or harder data leads to better results. |
| Approach: | They examine how well large language models generalize across different task difficulties by using a large dataset and a well-established difficulty metric. |
| Outcome: | The results show that training on hard data can't achieve consistent improvements across the full range of difficulties. |
Automated Concatenation of Embeddings for Structured Prediction (2021.acl-long)
Copied to clipboard
| Challenge: | Recent work shows that better word representations can be obtained by concatenating different types of embeddings. |
| Approach: | They propose to automate the process of finding better concatenated embeddings for structured prediction tasks by concatending different types of embeddables. |
| Outcome: | The proposed approach outperforms baselines and achieves state-of-the-art with fine-tuned embeddings on 6 tasks and 21 datasets. |
Multi-View Cross-Lingual Structured Prediction with Minimum Supervision (2021.acl-long)
Copied to clipboard
| Challenge: | Existing work on cross-lingual transfer learning focuses on transferring knowledge from high-resource languages to low-resourced ones. |
| Approach: | They propose a multi-view framework that integrates multiple source models into an aggregated source view and transfers it to a target view based on a task-specific model. |
| Outcome: | The proposed framework improves on three structured prediction tasks on 16 datasets. |
Investigating Context Faithfulness in Large Language Models: The Roles of Memory Strength and Evidence Style (2025.findings-acl)
Copied to clipboard
| Challenge: | Retrieval-augmented generation improves Large Language Models (LLMs) by integrating external information into the response generation process. |
| Approach: | They investigate the impact of memory strength and evidence presentation on LLMs’ receptiveness to external evidence by measuring the divergence in LLM responses to different paraphrases of the same question. |
| Outcome: | The proposed method improves Large Language Models (LLMs) by integrating external information into the response generation process. |
ITA: Image-Text Alignments for Multi-Modal Named Entity Recognition (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent work on Multi-modal Named Entity Recognition (MNER) relies on image information to model interactions between image and text representations. |
| Approach: | They propose to align image features into the textual space to better utilize attention mechanisms . they use regional object tags, captions and optical characters as visual contexts . |
| Outcome: | The proposed model can achieve state-of-the-art accuracy on multi-modal Named Entity Recognition datasets even without image information. |
Structure-Level Knowledge Distillation For Multilingual Sequence Labeling (2020.acl-main)
Copied to clipboard
| Challenge: | Existing multilingual models still underperform individual monolingual models due to model capacity limitations. |
| Approach: | They propose to distill the structural knowledge of several monolingual models (teachers) to the unified multilingual model (student). |
| Outcome: | The proposed model outperforms strong baseline models and teacher models on 4 multilingual tasks with 25 datasets and has stronger zero-shot generalizability. |
MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in entity retrieval ignore the property that meanings of entity mentions diverge in different contexts and are related to various portions of descriptions. |
| Approach: | They propose a novel approach that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method. |
| Outcome: | The proposed approach achieves state-of-the-art performance on ZESHEL and improves quality of candidates on three standard Entity Linking datasets. |
More Embeddings, Better Sequence Labelers? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work suggests contextual embeddings improve sequence labeling accuracy . but, there is no definite conclusion on whether concatenating different kinds of embeddables is effective . |
| Approach: | They propose a family of contextual embeddings that improves sequence labeling accuracy . they conduct extensive experiments on 3 tasks over 18 datasets and 8 languages . |
| Outcome: | The proposed family of contextual embeddings improves the accuracy of sequence labelers over non-contextual embedders. |
Answering Legal Questions by Learning Neural Attentive Text Representation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for retrieval-based question answering are limited by legal documents and long and complicated documents. |
| Approach: | They propose a retrieval-based model for answering legal questions at the article level by learning neural attentive text representation. |
| Outcome: | The proposed model outperforms state-of-the-art retrieval-based methods on an annotated corpus of 5,922 Vietnamese legal questions in terms of recall and NDCG. |
An Investigation of Potential Function Designs for Neural CRF (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to sequence labeling are based on the neural linear-chain CRF model. |
| Approach: | They propose a series of increasingly expressive potential functions for neural CRF models that integrate emission and transition functions and explicitly take contextual words as input. |
| Outcome: | The proposed model consistently achieves the best performance on the decomposed quadrilinear potential function based on the representations of two neighboring labels and two neighbored words. |
AIN: Fast and Accurate Sequence Labeling with Approximate Inference Network (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to sequence labeling require sequential computation that makes parallelization impossible. |
| Approach: | They propose to employ a parallelizable approximate variational inference algorithm for the CRF model. |
| Outcome: | The proposed approach improves decoding speed and accuracy with long sentences and is parallelizable for faster training and prediction. |