Papers by Tran Hoang
CareerPathKG: Knowledge Graph Integrated Framework for Career Intelligence (2026.eacl-industry)
Copied to clipboard
| Challenge: | a new framework for career orientation is needed to address the challenges of the labor market . a recent study found that traditional ML and large language models are brittle when faced with heterogeneous job descriptions . |
| Approach: | They propose a career-path knowledge graph-based recruitment framework to capture occupations, skill requirements and career transitions using standardized taxonomies enriched with job-posting data. |
| Outcome: | The proposed framework captures occupations, skill requirements, and career transitions using standardized taxonomies enriched with job-posting data. |
Diffusion Directed Acyclic Transformer for Non-Autoregressive Machine Translation (2025.acl-short)
Copied to clipboard
| Challenge: | Non-autoregressive transformers (NATs) often encounter performance challenges due to the multi-modality problem. |
| Approach: | They propose a direct-acyclic transformer (DAT) that captures multiple translation modalities to paths in a Directed Acyclic Graph (DAG) this allows the model to integrate latent variables into the model, which is crucial for DAT to achieve state-of-the-art performance. |
| Outcome: | The proposed model captures multiple translation modalities to paths in a Directed Acyclic Graph (DAG) but the collaboration with the latent variable introduced through the Glancing training is crucial for the model to attain state-of-the-art performance. |
PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | We present a high-quality and large-scale Vietnamese-English parallel dataset . our dataset is 2.9M pairs larger than the benchmark Vietnamese- English corpus . |
| Approach: | They present a large-scale Vietnamese-English parallel dataset with 3.02M sentence pairs . they compare strong neural baselines and well-known automatic translation engines . |
| Outcome: | The proposed dataset is 2.9M pairs larger than the benchmark Vietnamese-English corpus IWSLT15. |
Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension (2022.acl-long)
Copied to clipboard
Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer
| Challenge: | Existing QA datasets rarely distinguish fine-grained reading skills, such as the understanding of varying narrative elements. |
| Approach: | They propose to use FairytaleQA to generate 10,580 questions based on 278 children-friendly stories to assess model's fine-grained learning skills. |
| Outcome: | The proposed dataset consists of 10,580 questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations. |
VLUE: A New Benchmark and Multi-task Knowledge Transfer Learning for Vietnamese Natural Language Understanding (2024.findings-naacl)
Copied to clipboard
| Challenge: | a lack of standard evaluation metrics and benchmarks makes it difficult to identify strengths of Vietnamese NLP models. |
| Approach: | They propose to establish a standardized set of benchmarks for Vietnamese NLU . they propose to evaluate Vietnamese language understanding models using a pre-trained model . |
| Outcome: | The proposed model combines proficiency of a multilingual pre-trained model with Vietnamese linguistic knowledge. |
ViHOS: Hate Speech Spans Detection for Vietnamese (2023.eacl-main)
Copied to clipboard
| Challenge: | Increasing use of social networking sites can cause problems for human moderators to review tagged comments. |
| Approach: | They present a dataset that contains 26k spans on 11k comments and detailed annotation guidelines . they also provide definitions of hateful and offensive spans in Vietnamese comments . |
| Outcome: | The proposed dataset shows that it is difficult to detect specific types of spans in the dataset . the dataset is the first human-annotated corpus containing 26k spans on 11k comments . |
Class based Influence Functions for Error Detection (2023.acl-short)
Copied to clipboard
Thang Nguyen-Duc, Hoang Thanh-Tung, Quan Hung Tran, Dang Huu-Tien, Hieu Nguyen, Anh T. V. Dau, Nghi Bui
| Challenge: | Influence functions (IFs) are powerful tools for detecting anomalous examples in large scale datasets. |
| Approach: | They propose a method to explain the instability of IFs by leveraging class information to improve the stability of ifs. |
| Outcome: | The proposed method improves performance and stability while incurring no additional computational cost. |
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking (2026.acl-industry)
Copied to clipboard
| Challenge: | Existing methods struggle with semantic ambiguity, homonyms, and complex linguistic structures, often trading accuracy for efficiency. |
| Approach: | They propose a Vietnamese fact-checking framework that integrates SER and TVC to achieve 78.97% strict accuracy. |
| Outcome: | The proposed framework achieves state-of-the-art accuracy with 78.97% strict accuracy on ISE-DSC01 and 80.82% on ViWikiFC while maintaining competitive accuracy. |
ViHealthBERT: Pre-trained Language Models for Vietnamese in Health Text Mining (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent large-scale language models show remarkable achievements in key NLP tasks such as Question Answering and Text Summarization. |
| Approach: | They propose a domain-specific pre-trained Vietnamese language model that outperforms the general domain language models. |
| Outcome: | The proposed model outperforms the general domain language models in Vietnamese datasets while outperforming the general-domain language models. |