Papers by Thanh Vu
A Capsule Network-based Embedding Model for Knowledge Graph Completion and Search Personalization (N19-1)
Copied to clipboard
| Challenge: | Existing knowledge graphs with billions of triples are incomplete, i.e., missing a lot of valid triples. |
| Approach: | They propose to embed relationship triples into a capsule network using a convolution layer and multiple filters to generate feature maps. |
| Outcome: | The proposed model outperforms strong search personalization baselines on two benchmark datasets and outperformed previous state-of-the-art models on WN18RR and FB15k-237. |
BERTweet: A pre-trained language model for English Tweets (2020.emnlp-demos)
Copied to clipboard
| Challenge: | Experiments show that BERTweet outperforms strong baselines RoBERTa-base and XLM-R-base on three Tweet NLP tasks: Part-of-speech tagging, Named-entity recognition and text classification. |
| Approach: | They propose to train a pre-trained language model for English Tweets using the RoBERTa pre training procedure and use it to train the model. |
| Outcome: | Experiments show that the model outperforms baseline models on three Tweet NLP tasks: Part-of-speech tagging, Named-entity recognition and text classification. |
VnCoreNLP: A Vietnamese Natural Language Processing Toolkit (N18-5)
Copied to clipboard
| Challenge: | Using word segmenters and POS taggers, Vietnamese NLP pipelines are no longer considered SOTA models for Vietnamese. |
| Approach: | They propose a Java NLP annotation pipeline for Vietnamese that provides rich linguistic annotations. |
| Outcome: | The proposed toolkit provides rich linguistic annotations to facilitate research work on Vietnamese NLP. |
A Fast and Accurate Vietnamese Word Segmenter (L18-1)
Copied to clipboard
| Challenge: | Experimental results show that our approach outperforms previous state-of-the-art approaches in terms of accuracy and performance speed. |
| Approach: | They propose a method where rules are stored in an exception structure and new rules are only added to correct segmentation errors. |
| Outcome: | The proposed approach outperforms existing methods on Vietnamese treebank benchmarks. |
Mastering the Craft of Data Synthesis for CodeLLMs (2025.naacl-long)
Copied to clipboard
Meng Chen, Philip Arthur, Qianyu Feng, Cong Duy Vu Hoang, Yu-Heng Hong, Mahdi Kazemi Moghaddam, Omid Nezami, Duc Thien Nguyen, Gioacchino Tangari, Duy Vu, Thanh Vu, Mark Johnson, Krishnaram Kenthapadi, Don Dharmasiri, Long Duong, Yuan-Fang Li
| Challenge: | Large language models (LLMs) have shown impressive performance in code understanding and generation. |
| Approach: | They propose a systematic review of large language models and their taxonomy and propose specialized LLMs for code-related tasks. |
| Outcome: | The proposed models have shown to be highly effective in coding tasks. |
MedDCR: Learning to Design Agentic Workflows for Medical Coding (2026.findings-acl)
Copied to clipboard
| Challenge: | Medical coding is the process of translating unstructured clinical notes into standardized diagnostic and procedural codes. |
| Approach: | They propose a closed-loop framework that treats workflow design as a learning problem. |
| Outcome: | The proposed framework outperforms state-of-the-art workflows on benchmark datasets and produces interpretable, adaptable workflows that better reflect real coding practice. |