Papers by Thanh Vu

6 papers
A Capsule Network-based Embedding Model for Knowledge Graph Completion and Search Personalization (N19-1)

Copied to clipboard

Challenge: Existing knowledge graphs with billions of triples are incomplete, i.e., missing a lot of valid triples.
Approach: They propose to embed relationship triples into a capsule network using a convolution layer and multiple filters to generate feature maps.
Outcome: The proposed model outperforms strong search personalization baselines on two benchmark datasets and outperformed previous state-of-the-art models on WN18RR and FB15k-237.
BERTweet: A pre-trained language model for English Tweets (2020.emnlp-demos)

Copied to clipboard

Challenge: Experiments show that BERTweet outperforms strong baselines RoBERTa-base and XLM-R-base on three Tweet NLP tasks: Part-of-speech tagging, Named-entity recognition and text classification.
Approach: They propose to train a pre-trained language model for English Tweets using the RoBERTa pre training procedure and use it to train the model.
Outcome: Experiments show that the model outperforms baseline models on three Tweet NLP tasks: Part-of-speech tagging, Named-entity recognition and text classification.
VnCoreNLP: A Vietnamese Natural Language Processing Toolkit (N18-5)

Copied to clipboard

Challenge: Using word segmenters and POS taggers, Vietnamese NLP pipelines are no longer considered SOTA models for Vietnamese.
Approach: They propose a Java NLP annotation pipeline for Vietnamese that provides rich linguistic annotations.
Outcome: The proposed toolkit provides rich linguistic annotations to facilitate research work on Vietnamese NLP.
A Fast and Accurate Vietnamese Word Segmenter (L18-1)

Copied to clipboard

Challenge: Experimental results show that our approach outperforms previous state-of-the-art approaches in terms of accuracy and performance speed.
Approach: They propose a method where rules are stored in an exception structure and new rules are only added to correct segmentation errors.
Outcome: The proposed approach outperforms existing methods on Vietnamese treebank benchmarks.
Mastering the Craft of Data Synthesis for CodeLLMs (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive performance in code understanding and generation.
Approach: They propose a systematic review of large language models and their taxonomy and propose specialized LLMs for code-related tasks.
Outcome: The proposed models have shown to be highly effective in coding tasks.
MedDCR: Learning to Design Agentic Workflows for Medical Coding (2026.findings-acl)

Copied to clipboard

Challenge: Medical coding is the process of translating unstructured clinical notes into standardized diagnostic and procedural codes.
Approach: They propose a closed-loop framework that treats workflow design as a learning problem.
Outcome: The proposed framework outperforms state-of-the-art workflows on benchmark datasets and produces interpretable, adaptable workflows that better reflect real coding practice.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations