Papers by Hung-Ting Chen

7 papers
CaLMQA: Exploring culturally specific long-form question answering across 23 languages (2025.acl-long)

Copied to clipboard

Challenge: Despite rising global usage of large language models, their ability to generate *long-form* answers to *culturally specific* questions remains unexplored in many languages.
Approach: They perform the first study of textual multilingual long-form QA by creating a dataset of culturally specific questions across 23 different languages.
Outcome: The results show that the best models make critical surface-level errors for many languages and their understanding of diverse cultures.
MovieCORE: COgnitive REasoning in Movies (2025.emnlp-main)

Copied to clipboard

Challenge: MovieCORE is a video question answering dataset that focuses on surface-level comprehension.
Approach: They propose a video question-answer dataset that uses large language models as thought agents to generate and refine high-quality question-anchor pairs.
Outcome: The proposed model improves model reasoning capabilities post-training by 25% . the proposed model is based on a large language model and is scalable to a wide range of tasks .
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for embodied agents focus on directly executing instructions without considering whether objects can be manipulated.
Approach: They propose a benchmark that evaluates embodied agents in dynamic environments . they use plug-and-play module that augments existing planners with explicit affordance reasoning .
Outcome: The proposed benchmark evaluates embodied agents in dynamic environments with unpredictable affordances . ADAPT significantly improves robustness and task success across seen and unseen environments .
Continually Improving Extractive QA via Human Feedback (2023.emnlp-main)

Copied to clipboard

Challenge: a study of extractive question answering systems using human feedback shows promising potential for continual learning.
Approach: They study extractive question answering system by using user feedback to improve it . they design and deploy an iterative approach where users ask questions and provide feedback .
Outcome: The proposed model improves over time across different data regimes and domains . human user feedback is more affordable and abundant than annotations provided by trained experts .
OCID-Ref: A 3D Robotic Dataset With Embodied Language For Clutter Scene Grounding (2021.naacl-main)

Copied to clipboard

Challenge: Visual grounding (VG) is a crucial task in natural language processing, computer vision, and robotics.
Approach: They propose a visual grounding task with referring expressions of occluded objects in a OCID-Ref dataset with 2,300 scenes and a point cloud input.
Outcome: The proposed dataset shows that it can handle 2D and 3D signals but referring to occluded objects remains challenging for the modern visual grounding systems.
Rich Knowledge Sources Bring Complex Knowledge Conflicts: Recalibrating Models to Reflect Conflicting Evidence (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on question answering models relies on retrieved documents for provenance, but recent studies show that models can retain vast amounts of factual knowledge . retrieval-based generation approaches combine parametric knowledge sources with a large number of retrieved evidence documents, achieving state-of-the-art performance on open retrieval datasets.
Approach: They propose to use parametric and parametric knowledge to generate free-form questions from retrieved evidence documents.
Outcome: The proposed model can use parametric and parametric knowledge to generate free-form answers from retrieved evidence documents.
Open-World Evaluation for Retrieving Diverse Perspectives (2025.naacl-long)

Copied to clipboard

Challenge: Existing retrieval systems only cover diverse perspectives on 33.74% of the examples . existing systems only focus on relevance to the question, ignoring diversity.
Approach: They build a Benchmark for Retrieval Diversity for Subjective questions (BERDS) based on a question and diverse perspectives associated with the question . they evaluate retrievers paired with a corpus to determine whether each document contains a perspective .
Outcome: The proposed approach improves retrieval diversity on complex questions . existing retrieval systems only cover diverse perspectives on 33.74% of the examples .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations