Papers by Chia-Chien Hung

8 papers
Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies show that incorporating demographic factors in language representations improves performance on downstream NLP tasks.
Approach: They use continuous language modeling and dynamic multi-task learning to adapt pre-trained Transformers to incorporate demographic information into their representations.
Outcome: The proposed model shows that the results are consistent with previous studies.
Linking Surface Facts to Large-Scale Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Open Information Extraction (OIE) methods extract facts in the form of triples . ambiguity of these triples hinders their downstream usage .
Approach: They propose a benchmark that measures fact linking performance on a granular triple slot level . they propose to use a system that can detect out-of-KG entities and predicates .
Outcome: The proposed benchmark can measure fact linking performance on a granular triple slot level while also measuring if a system can recognize that a surface form has no match in the existing KG.
Multi2WOZ: A Robust Multilingual Dataset and Conversational Pretraining for Task-Oriented Dialog (2022.naacl-main)

Copied to clipboard

Challenge: Task-oriented dialog (TOD) is arguably one of the most popular natural language processing (NLP) application areas.
Approach: They propose a multilingual multi-domain TOD dataset that spans four languages . they use a framework for multilingual conversational specialization of pretrained language models .
Outcome: The proposed datasets show that they perform better than existing datasets in English . the proposed framework allows for sample-efficient few-shot transfer for TOD tasks .
On Synthesizing Data for Context Attribution in Question Answering (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a tendency to hallucinate, resulting in false or misleading answers.
Approach: They propose a novel generative strategy for synthesizing context attribution data.
Outcome: The proposed approach is highly effective for fine-tuning small LMs for context attribution in different QA tasks and domains.
MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to DDx are limited by single-dataset evaluations, isolated optimization of components, unrealistic assumptions about complete patient profiles, and single-attempt diagnosis.
Approach: They propose a Modular Explainable DDx Agent framework that allows physicians to iteratively refine a ranked list of possible diseases based on symptoms, antecedents, and medical knowledge.
Outcome: The proposed framework achieves over 10% accuracy improvements in interactive DDx across large and small LLMs while offering critical explainability into its diagnostic reasoning process.
ANHALTEN: Cross-Lingual Transfer for German Token-Level Reference-Free Hallucination Detection (2024.acl-srw)

Copied to clipboard

Challenge: ANHALTEN is a new evaluation dataset that extends the English hallucination detection dataset to German.
Approach: They propose a dataset that extends the English hallucination detection dataset to German . they show that larger context length leads to better halluciation detection in german .
Outcome: ANHALTEN is the first evaluation dataset that extends the English hallucination detection dataset to German.
TADA: Efficient Task-Agnostic Domain Adaptation for Transformers (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained transformer-based language models are limited in their expressiveness and domain knowledge.
Approach: They propose a task-agnostic domain adaptation method which is modular, parameter-efficient, and data-efficient.
Outcome: The proposed method is efficient and modular, parameter-efficient, and data-efficient.
DS-TOD: Efficient Domain Specialization for Task-Oriented Dialog (2022.findings-acl)

Copied to clipboard

Challenge: Recent work shows that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining.
Approach: They propose a resource-efficient and modular domain specialization by means of domain adapters in which domain knowledge is encoded.
Outcome: The proposed framework extracts domain-specific terms and then uses them to build DomainCC and DomainReddit resources based on masked language modeling and response selection objectives.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations