Papers by Chenhao Tan

33 papers
Literature Meets Data: A Synergistic Approach to Hypothesis Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for hypothesis generation are theory-driven and data-driven, but they lack the computational power to complement each other.
Approach: They develop a method that combines literature-based insights with data to perform LLM-powered hypothesis generation.
Outcome: The proposed method outperforms baseline methods on five datasets and shows human accuracy improves on deception detection and AI generated content detection tasks.
CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support (2024.findings-acl)

Copied to clipboard

Challenge: Literature review requires researchers to synthesize a large amount of information.
Approach: They propose to use LLMs to generate hierarchical organizations from a set of studies . they use a human-in-the-loop process to correct errors in LLM-generated hierarchies .
Outcome: The proposed model improves assignment of studies to categories by 12.6 F1 points.
Explaining Why: How Instructions and User Interfaces Impact Annotator Rationales When Labeling Text Data (2022.naacl-main)

Copied to clipboard

Challenge: In the context of data labeling, researchers are interested in having humans select rationales .
Approach: They conducted an online user study to understand how humans select rationales . they found that participants were near unanimous in their data labels .
Outcome: The results show that participants selected 12% of input tokens as rationales, but fewer if unable to drag over multiple tokens at once.
What Gets Echoed? Understanding the “Pointers” in Explanations of Persuasive Arguments (D19-1)

Copied to clipboard

Challenge: Explanations are central to everyday life, and are a topic of growing interest in the AI community.
Approach: They propose a word-level prediction task to investigate how explanations selectively reuse information from what is being explained.
Outcome: The proposed features have strong predictive power on the echoing of a word in an explanation, and enhance neural methods of generating explanations.
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing frameworks for LLM-as-a-judge use zero-shot setting without consulting any human input, which leads to low alignment, or fine-tune LLMs on labeled data, which requires a non-trivial number of samples.
Approach: They propose a hypothesis-guided evaluation framework that uses a small corpus of human evaluations to generate more detailed rubrics for human judgments and incorporates a checklist-like approach to combine LLM’s assigned scores on each decomposed dimension to acquire overall scores.
Outcome: The proposed framework outperforms existing frameworks in both human rankings and human scores with 30 human evaluations and fine-tunes LLMs on labeled data with 3 times more human evaluation by 11.95%.
Language of Bargaining (2023.acl-long)

Copied to clipboard

Challenge: a new dataset is being developed to study how language shapes bilateral bargaining . a recent study examined the use of language in negotiation education .
Approach: They propose a dataset to study how language shapes bilateral bargaining . they recruit participants via behavioral labs instead of crowdsourcing platforms .
Outcome: The proposed dataset is based on an exercise in negotiation education . it shows that when subjects can talk, negotiations finish faster and prices drop .
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge (2026.acl-demo)

Copied to clipboard

Challenge: AutoChecklist is an open-source library that unifies checklist-based evaluation into composable pipelines.
Approach: They propose an open-source library that unifies checklist-based evaluation into composable pipelines.
Outcome: The open-source library unifies checklist-based evaluation into composable pipelines.
CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing metrics lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports.
Approach: They propose a tabular framework with E**xpert-curated labels and an attribute-level comparison for radiology report evaluation (**CLEAR)
Outcome: The proposed framework can extract clinical attributes and provide automated metrics that are strongly aligned with clinical judgment.
The Impossibility of Fair LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing frameworks for evaluating large language models do not extend to general-purpose AI contexts or are infeasible in practice.
Approach: They analyze a variety of technical fairness frameworks to find inherent challenges . they find that each framework does not logically extend to the general-purpose AI context .
Outcome: The proposed frameworks do not logically extend to the general-purpose AI context or are infeasible in practice due to large amounts of unstructured training data and potential combinations of human populations, use cases, and sensitive attributes.
When Internalization Fails: Finding Better Targets for Reasoning Compression (2026.findings-acl)

Copied to clipboard

Challenge: Reasoning language models generate long reasoning traces that increase latency and cost.
Approach: They compare three approaches to shorten reasoning traces by inference-time truncation . they use Implicit Chain-of-Thought-style curricula that progressively shorten the teacher trace .
Outcome: The proposed methods work well on GSM8K and multiplication tasks.
FLamE: Few-shot Learning from Natural Language Explanations (2023.acl-long)

Copied to clipboard

Challenge: Recent work has shown limited utility of natural language explanations in improving classification.
Approach: They propose a two-stage few-shot learning framework that generates explanations and fine-tunes a smaller model with generated explanations.
Outcome: The proposed framework increases inference accuracy over strong baselines, but human evaluation reveals that the majority of generated explanations does not adequately justify classification decisions.
Personalized Benchmarking: Evaluating LLMs by Individual Preferences (2026.findings-acl)

Copied to clipboard

Challenge: Current benchmarks average preferences across all users to compute aggregate ratings . this overlooks individual user preferences when establishing model rankings .
Approach: They compute personalized model rankings using ELO ratings and Bradley-Terry coefficients . they find users exhibit substantial heterogeneity in topical interests and communication styles .
Outcome: The results show that individual rankings of LLM models diverge dramatically from aggregate rankings . a compact combination of topic and style features provides a useful feature space .
CPsyExam: A Chinese Benchmark for Evaluating Psychology using Examinations (2025.coling-main)

Copied to clipboard

Challenge: CPsyExam prioritizes psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios.
Approach: They propose a psychological benchmark, CPsyExam, constructed from questions from Chinese examination systems.
Outcome: The proposed benchmark prioritizes psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios.
Human-Centered Evaluation of Explanations (2022.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide an overview of human-centered evaluations of explanations .
Approach: This tutorial will provide an overview of human-centered evaluations of explanations . it will introduce the psychological foundation of explanation and types of NLP explanations.
Outcome: This tutorial will provide an overview of human-centered evaluations of explanations . it will cover the two categories of evaluation: evaluation based on human-annotated explanations and evaluation with human-subjects studies.
Explanation in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events.
Approach: They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations.
Outcome: The proposed methods are based on the models of large language models (LLMs) and their opaque nature.
CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets lack consulting knowledge, resulting in LLMs lacking professional consulting competence.
Approach: They propose a report-based multi-turn dialogue reconstruction framework for Chinese psychological counseling that uses large language models to assist counseling.
Outcome: The proposed framework is open-source and can be used in future research.
MoVa: Towards Generalizable Classification of Human Morals and Values (2025.emnlp-main)

Copied to clipboard

Challenge: Identifying human morals and values embedded in language is essential to empirical studies of communication.
Approach: They propose a framework for generalizable classification of human morals and values . they recommend a classification strategy that scores all related concepts simultaneously .
Outcome: The proposed method outperforms fine-tuned models across domains and frameworks.
Entity-Based Evaluation of Political Bias in Automatic Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have shown that NLP systems may encode social biases, but the *political* bias of summarization models remains relatively unknown.
Approach: They use an entity replacement method to examine the portrayal of politicians in automatically generated summaries.
Outcome: The proposed model can control for the content of the source document and can be used to predict the ideal quality of summarization models.
What to Learn, and How: Toward Effective Learning from Rationales (2022.findings-acl)

Copied to clipboard

Challenge: Increasing interest in learning from rationales has led to the use of human-annotated explanations to inject useful inductive biases into models.
Approach: They propose several novel loss functions and learning strategies to exploit human rationales to augment model prediction accuracy.
Outcome: The proposed learning strategies improve on three datasets with human rationales and show that they are more efficient than baselines.
On Positivity Bias in Negative Reviews (2021.acl-short)

Copied to clipboard

Challenge: Existing studies have shown positive words are more frequently used in negative reviews . however, it remains unclear whether the Pollyanna hypothesis holds in negative review .
Approach: They validate the Pollyanna hypothesis that positive words occur more frequently than negative words in human expressions . they use a variety of review datasets to examine the use of positive and negative words .
Outcome: The results confirm the pollyanna hypothesis that positive words occur more frequently than negative words in human expressions.
GPT-4V Cannot Generate Radiology Reports Yet (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) are becoming multimodal, and GPT-4 models are supposed to possess advanced skills across a wide range of domains, including high-stakes scenarios such as medicine.
Approach: They perform a systematic evaluation of GPT-4 in generating radiology reports across three chest X-ray report benchmarks: MIMIC-CXR, CheXpert Plus, and IU X ray.
Outcome: The proposed model fails in lexical and clinical efficacy metrics . the distributions of model-predicted labels remain constant regardless of groundtruth conditions on the image, suggesting that the model is not interpreting chest X-rays meaningfully.
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions (2025.findings-naacl)

Copied to clipboard

Challenge: CaseSumm is a dataset for long-context summarization in the legal domain . human groundtruth summaries are often not available for legal summarizing .
Approach: They propose a dataset for long-context summarization that includes SCOTUS opinions and their official summaries.
Outcome: The proposed dataset is the largest open legal case summarization dataset . it outperforms larger models on automatic metrics and human evaluation .
Active Example Selection for In-Context Learning (2022.emnlp-main)

Copied to clipboard

Challenge: In-context learning performance is unstable across samples of examples, suggesting the idiosyncrasies of how language models acquire information.
Approach: They propose a reinforcement learning algorithm for identifying generalizable policies to select demonstration examples and propose 'in-context learning' performance can be highly unstable across samples of examples, suggesting the idiosyncrasies of how language models acquire information.
Outcome: The proposed model can perform tasks with examples with a 5.8% improvement on GPT-2 and GPT-3, but the improvement diminishes on larger models, suggesting emerging capabilities of large language models.
Evaluating and Characterizing Human Rationales (2020.emnlp-main)

Copied to clipboard

Challenge: a new study examines how human rationales perform on automatic metrics . human-generated rationale evaluation is difficult because of its ambiguity .
Approach: They propose to use model-dependent baseline performance to evaluate rationale quality . they propose to also use "fidelity curves" to reveal properties such as irrelevance and redundancy .
Outcome: The proposed methods characterize rationale quality based on model retraining and using "fidelity curves" the proposed methods lead to actionable suggestions for evaluating and characterizing rationales .
Learning to Ignore Adversarial Attacks (2023.eacl-main)

Copied to clipboard

Challenge: Despite the strong performance of current NLP models, they can be brittle against adversarial inputs.
Approach: They propose a rationale model that explicitly learns to ignore adversarial tokens . their approach leads to sizable improvements in robustness over baseline models .
Outcome: The proposed model outperforms data augmentation with adversarial examples and closes the gap between model performance and an attacked test set.
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing automated metrics fail to align with real-world physician preferences.
Approach: They propose a pipeline that distills real user feedback into structured checklists for note evaluation that are interpretable, grounded in human feedback, and enforceable by LLM-based evaluators.
Outcome: The proposed checklist outperforms baseline evaluations in coverage, diversity, and predictive power for human ratings.
On the Diversity and Limits of Human Explanations (2022.naacl-main)

Copied to clipboard

Challenge: a growing effort in NLP aims to build datasets of human explanations, but it remains unclear whether they serve their intended goals.
Approach: They argue that the term "explanation" is overloaded and refers to a broad range of notions with different properties and ramifications.
Outcome: The proposed datasets examine the diversity of explanations and their use in NLP.
Neural Models for Documents with Metadata (P18-1)

Copied to clipboard

Challenge: specialized models are often used to model text corpora without metadata . specialized algorithms are not widely used in the digital humanities and political science fields .
Approach: They propose a general neural framework based on topic models to enable customization of metadata.
Outcome: The proposed framework achieves strong performance with a manageable tradeoff between perplexity, coherence, and sparsity.
Many Faces of Feature Importance: Comparing Built-in and Post-hoc Feature Importance in Text Classification (D19-1)

Copied to clipboard

Challenge: Feature importance is commonly used to explain machine predictions . however, the consistency of feature importance via different methods remains understudied .
Approach: They compare feature importance from built-in mechanisms and post-hoc methods that approximate model behavior to find similarities between models.
Outcome: The proposed methods show that features from traditional models are more similar with each other than with deep learning models.
Decision-Focused Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing summarization methods define relevance based on textual information alone without incorporating insights about a particular decision.
Approach: They propose a method that summarizes relevant information for a decision using full text . they then build a model that makes the decision based on the full text while accounting for textual non-redundancy.
Outcome: The proposed method outperforms text-only summarization methods and model-based explanation methods in decision faithfulness and representativeness.
Characterizing the Value of Information in Medical Notes (2020.findings-emnlp)

Copied to clipboard

Challenge: Obtaining and analyzing information is critical for the diagnosis, prognosis, treatment, and prevention of disease.
Approach: They propose a probing framework to select parts of notes that enable more accurate predictions than using all notes.
Outcome: The proposed framework achieves better predictive performance with only 6.8% of all tokens for readmission prediction.
Ecologically Valid Explanations for Label Variation in NLI (2023.findings-emnlp)

Copied to clipboard

Challenge: Human label variation exists in many natural language processing tasks, including NLI .
Approach: They build an English dataset of 1,415 ecologically valid explanations for 122 MNLI items . they find that people can systematically vary on their interpretation .
Outcome: The proposed dataset contains 1,415 ecologically valid explanations for 122 items . the results show that people can vary on interpretation and highlight differences .
No Permanent Friends or Enemies: Tracking Relationships between Nations from News (N19-1)

Copied to clipboard

Challenge: Understanding complex international relations is important but challenging for civilians . topic models and neural models have been proposed to explore relations without supervision .
Approach: They propose an unsupervised neural model that integrates linguistic insights into the model to infer relations between nations from news articles.
Outcome: The proposed model outperforms baselines from topic models and hidden Markov models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations