Papers by Thuy Vu

9 papers
CDA: a Cost Efficient Content-based Multilingual Web Document Aligner (2021.eacl-main)

Copied to clipboard

Challenge: a Content-based document alignment approach is an efficient way to align multilingual web documents based on content.
Approach: They propose a Content-based document alignment approach to align multilingual web documents based on content in parallel training data for machine translation systems.
Outcome: The proposed method achieves comparable performance with state-of-the-art systems in the WMT-16 Bilingual Document Alignment Shared Task benchmark while operating in multilingual space.
Question-Answer Sentence Graph for Joint Modeling Answer Selection (2023.eacl-main)

Copied to clipboard

Challenge: Existing approaches to automate Question Answering (QA) are graph-based and can target large text databases.
Approach: They propose graph-based approaches for Answer Sentence Selection (AS2) . they train and integrate state-of-the-art (SOTA) models for computing scores .
Outcome: The proposed approach outperforms baseline models on academic benchmarks and a real-world dataset on unseen queries.
Joint Models for Answer Verification in Question Answering Systems (2021.acl-long)

Copied to clipboard

Challenge: Using a joint approach, we found that the model is more efficient than those developed in machine reading (MR) work.
Approach: They propose a joint model for selecting correct answer sentences among the top k provided by answer sentence selection modules.
Outcome: The proposed model improves on WikiQA, TREC-QA, and a real-world dataset.
Double Retrieval and Ranking for Accurate Question Answering (2023.findings-eacl)

Copied to clipboard

Challenge: Recent work shows that answer verification models can improve the state of the art in Question Answering . despite the fact that the supporting candidates are ranked only according to the relevancy with the question, the model still lacks the support needed for other answer candidates.
Approach: They propose a double reranking model that selects the best support for each target answer . they propose 'second neural retrieval stage' to encode question and answer pair as query .
Outcome: The proposed approach improves the state of the art in Question Answering . the proposed model ranked candidates according to relevancy and not the answer . but the proposed approach fails to provide the best support .
AVA: an Automatic eValuation Approach for Question Answering Systems (2021.naacl-main)

Copied to clipboard

Challenge: AVA is an automatic evaluation approach for question answering . it uses transformer-based language models to encode question, answer, and reference texts .
Approach: They propose an automatic evaluation approach for Question Answering that uses Transformer-based language models to encode question, answer, and reference texts.
Outcome: AVA can estimate system Accuracy with an error lower than 7% at 95% confidence level . the proposed approach achieves 74.7% F1 score in predicting human judgment for single answers .
FocusQA: Open-Domain Question Answering with a Context in Focus (2022.findings-emnlp)

Copied to clipboard

Challenge: a new method for question answering with a context in focus simulates a free interaction with QA systems.
Approach: They introduce question answering with a cotext in focus task that simulates a free interaction with QA systems.
Outcome: The proposed model outperforms state-of-the-art models for question answering with a context in focus up to 21.3% absolute points.
Retrieving Support to Rank Answers in Open-Domain Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: a novel question answering architecture retrieves content relevant to the combined pair . previous work on automatic claim verification has shown hallucinations .
Approach: They propose a question-answer architecture that prioritizes supporting evidence . it retrieves paragraphs that directly substantiate the correctness of a with respect to q .
Outcome: The proposed approach can be used by large language models to retrieve explanatory paragraphs that ground their reasoning.
Reference-based Weak Supervision for Answer Sentence Selection using Web Data (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing solutions for QA use weakly-supervised data from Web.
Approach: They propose a data pipeline that harvests weakly-supervised answer sentences from Web data . they train TANDA models, which are the state of the art for AS2 .
Outcome: The proposed pipeline improves on three different datasets and sets the state-of-the-art models to P@1=90.1% and MAP=92.9%.
CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic Data (2026.acl-long)

Copied to clipboard

Challenge: Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress.
Approach: They propose a knowledge Graph retrieval tool that can translate natural language questions into structured queries.
Outcome: Extensive experiments show that CypherSmith achieves state-of-the-art performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations