Papers by Chenhao Tan
Literature Meets Data: A Synergistic Approach to Hypothesis Generation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for hypothesis generation are theory-driven and data-driven, but they lack the computational power to complement each other. |
| Approach: | They develop a method that combines literature-based insights with data to perform LLM-powered hypothesis generation. |
| Outcome: | The proposed method outperforms baseline methods on five datasets and shows human accuracy improves on deception detection and AI generated content detection tasks. |
CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support (2024.findings-acl)
Copied to clipboard
Chao-Chun Hsu, Erin Bransom, Jenna Sparks, Bailey Kuehl, Chenhao Tan, David Wadden, Lucy Wang, Aakanksha Naik
| Challenge: | Literature review requires researchers to synthesize a large amount of information. |
| Approach: | They propose to use LLMs to generate hierarchical organizations from a set of studies . they use a human-in-the-loop process to correct errors in LLM-generated hierarchies . |
| Outcome: | The proposed model improves assignment of studies to categories by 12.6 F1 points. |
Explaining Why: How Instructions and User Interfaces Impact Annotator Rationales When Labeling Text Data (2022.naacl-main)
Copied to clipboard
Jamar Sullivan Jr., Will Brackenbury, Andrew McNutt, Kevin Bryson, Kwam Byll, Yuxin Chen, Michael Littman, Chenhao Tan, Blase Ur
| Challenge: | In the context of data labeling, researchers are interested in having humans select rationales . |
| Approach: | They conducted an online user study to understand how humans select rationales . they found that participants were near unanimous in their data labels . |
| Outcome: | The results show that participants selected 12% of input tokens as rationales, but fewer if unable to drag over multiple tokens at once. |
What Gets Echoed? Understanding the “Pointers” in Explanations of Persuasive Arguments (D19-1)
Copied to clipboard
| Challenge: | Explanations are central to everyday life, and are a topic of growing interest in the AI community. |
| Approach: | They propose a word-level prediction task to investigate how explanations selectively reuse information from what is being explained. |
| Outcome: | The proposed features have strong predictive power on the echoing of a word in an explanation, and enhance neural methods of generating explanations. |
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing frameworks for LLM-as-a-judge use zero-shot setting without consulting any human input, which leads to low alignment, or fine-tune LLMs on labeled data, which requires a non-trivial number of samples. |
| Approach: | They propose a hypothesis-guided evaluation framework that uses a small corpus of human evaluations to generate more detailed rubrics for human judgments and incorporates a checklist-like approach to combine LLM’s assigned scores on each decomposed dimension to acquire overall scores. |
| Outcome: | The proposed framework outperforms existing frameworks in both human rankings and human scores with 30 human evaluations and fine-tunes LLMs on labeled data with 3 times more human evaluation by 11.95%. |
Language of Bargaining (2023.acl-long)
Copied to clipboard
| Challenge: | a new dataset is being developed to study how language shapes bilateral bargaining . a recent study examined the use of language in negotiation education . |
| Approach: | They propose a dataset to study how language shapes bilateral bargaining . they recruit participants via behavioral labs instead of crowdsourcing platforms . |
| Outcome: | The proposed dataset is based on an exercise in negotiation education . it shows that when subjects can talk, negotiations finish faster and prices drop . |
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge (2026.acl-demo)
Copied to clipboard
| Challenge: | AutoChecklist is an open-source library that unifies checklist-based evaluation into composable pipelines. |
| Approach: | They propose an open-source library that unifies checklist-based evaluation into composable pipelines. |
| Outcome: | The open-source library unifies checklist-based evaluation into composable pipelines. |
CLEAR: A Clinically Grounded Tabular Framework for Radiology Report Evaluation (2025.findings-emnlp)
Copied to clipboard
Yuyang Jiang, Chacha Chen, Shengyuan Wang, Feng Li, Zecong Tang, Benjamin M. Mervak, Lydia Chelala, Christopher M Straus, Reve Chahine, Samuel G. Armato Iii, Chenhao Tan
| Challenge: | Existing metrics lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports. |
| Approach: | They propose a tabular framework with E**xpert-curated labels and an attribute-level comparison for radiology report evaluation (**CLEAR) |
| Outcome: | The proposed framework can extract clinical attributes and provide automated metrics that are strongly aligned with clinical judgment. |
The Impossibility of Fair LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Existing frameworks for evaluating large language models do not extend to general-purpose AI contexts or are infeasible in practice. |
| Approach: | They analyze a variety of technical fairness frameworks to find inherent challenges . they find that each framework does not logically extend to the general-purpose AI context . |
| Outcome: | The proposed frameworks do not logically extend to the general-purpose AI context or are infeasible in practice due to large amounts of unstructured training data and potential combinations of human populations, use cases, and sensitive attributes. |
When Internalization Fails: Finding Better Targets for Reasoning Compression (2026.findings-acl)
Copied to clipboard
| Challenge: | Reasoning language models generate long reasoning traces that increase latency and cost. |
| Approach: | They compare three approaches to shorten reasoning traces by inference-time truncation . they use Implicit Chain-of-Thought-style curricula that progressively shorten the teacher trace . |
| Outcome: | The proposed methods work well on GSM8K and multiplication tasks. |
FLamE: Few-shot Learning from Natural Language Explanations (2023.acl-long)
Copied to clipboard
| Challenge: | Recent work has shown limited utility of natural language explanations in improving classification. |
| Approach: | They propose a two-stage few-shot learning framework that generates explanations and fine-tunes a smaller model with generated explanations. |
| Outcome: | The proposed framework increases inference accuracy over strong baselines, but human evaluation reveals that the majority of generated explanations does not adequately justify classification decisions. |
Personalized Benchmarking: Evaluating LLMs by Individual Preferences (2026.findings-acl)
Copied to clipboard
| Challenge: | Current benchmarks average preferences across all users to compute aggregate ratings . this overlooks individual user preferences when establishing model rankings . |
| Approach: | They compute personalized model rankings using ELO ratings and Bradley-Terry coefficients . they find users exhibit substantial heterogeneity in topical interests and communication styles . |
| Outcome: | The results show that individual rankings of LLM models diverge dramatically from aggregate rankings . a compact combination of topic and style features provides a useful feature space . |
CPsyExam: A Chinese Benchmark for Evaluating Psychology using Examinations (2025.coling-main)
Copied to clipboard
Jiahao Zhao, Jingwei Zhu, Minghuan Tan, Min Yang, Renhao Li, Yang Di, Chenhao Zhang, Guancheng Ye, Chengming Li, Xiping Hu, Derek F. Wong
| Challenge: | CPsyExam prioritizes psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios. |
| Approach: | They propose a psychological benchmark, CPsyExam, constructed from questions from Chinese examination systems. |
| Outcome: | The proposed benchmark prioritizes psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios. |
Human-Centered Evaluation of Explanations (2022.naacl-tutorials)
Copied to clipboard
Jordan Boyd-Graber, Samuel Carton, Shi Feng, Q. Vera Liao, Tania Lombrozo, Alison Smith-Renner, Chenhao Tan
| Challenge: | This tutorial will provide an overview of human-centered evaluations of explanations . |
| Approach: | This tutorial will provide an overview of human-centered evaluations of explanations . it will introduce the psychological foundation of explanation and types of NLP explanations. |
| Outcome: | This tutorial will provide an overview of human-centered evaluations of explanations . it will cover the two categories of evaluation: evaluation based on human-annotated explanations and evaluation with human-subjects studies. |
Explanation in the Era of Large Language Models (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | Explanation has long been a part of communication, where humans use language to elucidate each other and transmit information about mechanisms of events. |
| Approach: | They review the opportunities and challenges of explanations in the era of large language models and examine how they can be used to generate explanations. |
| Outcome: | The proposed methods are based on the models of large language models (LLMs) and their opaque nature. |
CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling (2024.findings-acl)
Copied to clipboard
Chenhao Zhang, Renhao Li, Minghuan Tan, Min Yang, Jingwei Zhu, Di Yang, Jiahao Zhao, Guancheng Ye, Chengming Li, Xiping Hu
| Challenge: | Existing datasets lack consulting knowledge, resulting in LLMs lacking professional consulting competence. |
| Approach: | They propose a report-based multi-turn dialogue reconstruction framework for Chinese psychological counseling that uses large language models to assist counseling. |
| Outcome: | The proposed framework is open-source and can be used in future research. |
MoVa: Towards Generalizable Classification of Human Morals and Values (2025.emnlp-main)
Copied to clipboard
Ziyu Chen, Junfei Sun, Chenxi Li, Tuan Dung Nguyen, Jing Yao, Xiaoyuan Yi, Xing Xie, Chenhao Tan, Lexing Xie
| Challenge: | Identifying human morals and values embedded in language is essential to empirical studies of communication. |
| Approach: | They propose a framework for generalizable classification of human morals and values . they recommend a classification strategy that scores all related concepts simultaneously . |
| Outcome: | The proposed method outperforms fine-tuned models across domains and frameworks. |
Entity-Based Evaluation of Political Bias in Automatic Summarization (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have shown that NLP systems may encode social biases, but the *political* bias of summarization models remains relatively unknown. |
| Approach: | They use an entity replacement method to examine the portrayal of politicians in automatically generated summaries. |
| Outcome: | The proposed model can control for the content of the source document and can be used to predict the ideal quality of summarization models. |
What to Learn, and How: Toward Effective Learning from Rationales (2022.findings-acl)
Copied to clipboard
| Challenge: | Increasing interest in learning from rationales has led to the use of human-annotated explanations to inject useful inductive biases into models. |
| Approach: | They propose several novel loss functions and learning strategies to exploit human rationales to augment model prediction accuracy. |
| Outcome: | The proposed learning strategies improve on three datasets with human rationales and show that they are more efficient than baselines. |
On Positivity Bias in Negative Reviews (2021.acl-short)
Copied to clipboard
| Challenge: | Existing studies have shown positive words are more frequently used in negative reviews . however, it remains unclear whether the Pollyanna hypothesis holds in negative review . |
| Approach: | They validate the Pollyanna hypothesis that positive words occur more frequently than negative words in human expressions . they use a variety of review datasets to examine the use of positive and negative words . |
| Outcome: | The results confirm the pollyanna hypothesis that positive words occur more frequently than negative words in human expressions. |
GPT-4V Cannot Generate Radiology Reports Yet (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are becoming multimodal, and GPT-4 models are supposed to possess advanced skills across a wide range of domains, including high-stakes scenarios such as medicine. |
| Approach: | They perform a systematic evaluation of GPT-4 in generating radiology reports across three chest X-ray report benchmarks: MIMIC-CXR, CheXpert Plus, and IU X ray. |
| Outcome: | The proposed model fails in lexical and clinical efficacy metrics . the distributions of model-predicted labels remain constant regardless of groundtruth conditions on the image, suggesting that the model is not interpreting chest X-rays meaningfully. |
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions (2025.findings-naacl)
Copied to clipboard
| Challenge: | CaseSumm is a dataset for long-context summarization in the legal domain . human groundtruth summaries are often not available for legal summarizing . |
| Approach: | They propose a dataset for long-context summarization that includes SCOTUS opinions and their official summaries. |
| Outcome: | The proposed dataset is the largest open legal case summarization dataset . it outperforms larger models on automatic metrics and human evaluation . |
Active Example Selection for In-Context Learning (2022.emnlp-main)
Copied to clipboard
| Challenge: | In-context learning performance is unstable across samples of examples, suggesting the idiosyncrasies of how language models acquire information. |
| Approach: | They propose a reinforcement learning algorithm for identifying generalizable policies to select demonstration examples and propose 'in-context learning' performance can be highly unstable across samples of examples, suggesting the idiosyncrasies of how language models acquire information. |
| Outcome: | The proposed model can perform tasks with examples with a 5.8% improvement on GPT-2 and GPT-3, but the improvement diminishes on larger models, suggesting emerging capabilities of large language models. |
Evaluating and Characterizing Human Rationales (2020.emnlp-main)
Copied to clipboard
| Challenge: | a new study examines how human rationales perform on automatic metrics . human-generated rationale evaluation is difficult because of its ambiguity . |
| Approach: | They propose to use model-dependent baseline performance to evaluate rationale quality . they propose to also use "fidelity curves" to reveal properties such as irrelevance and redundancy . |
| Outcome: | The proposed methods characterize rationale quality based on model retraining and using "fidelity curves" the proposed methods lead to actionable suggestions for evaluating and characterizing rationales . |
Learning to Ignore Adversarial Attacks (2023.eacl-main)
Copied to clipboard
| Challenge: | Despite the strong performance of current NLP models, they can be brittle against adversarial inputs. |
| Approach: | They propose a rationale model that explicitly learns to ignore adversarial tokens . their approach leads to sizable improvements in robustness over baseline models . |
| Outcome: | The proposed model outperforms data augmentation with adversarial examples and closes the gap between model performance and an attacked test set. |
From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing automated metrics fail to align with real-world physician preferences. |
| Approach: | They propose a pipeline that distills real user feedback into structured checklists for note evaluation that are interpretable, grounded in human feedback, and enforceable by LLM-based evaluators. |
| Outcome: | The proposed checklist outperforms baseline evaluations in coverage, diversity, and predictive power for human ratings. |
On the Diversity and Limits of Human Explanations (2022.naacl-main)
Copied to clipboard
| Challenge: | a growing effort in NLP aims to build datasets of human explanations, but it remains unclear whether they serve their intended goals. |
| Approach: | They argue that the term "explanation" is overloaded and refers to a broad range of notions with different properties and ramifications. |
| Outcome: | The proposed datasets examine the diversity of explanations and their use in NLP. |
Neural Models for Documents with Metadata (P18-1)
Copied to clipboard
| Challenge: | specialized models are often used to model text corpora without metadata . specialized algorithms are not widely used in the digital humanities and political science fields . |
| Approach: | They propose a general neural framework based on topic models to enable customization of metadata. |
| Outcome: | The proposed framework achieves strong performance with a manageable tradeoff between perplexity, coherence, and sparsity. |
Many Faces of Feature Importance: Comparing Built-in and Post-hoc Feature Importance in Text Classification (D19-1)
Copied to clipboard
| Challenge: | Feature importance is commonly used to explain machine predictions . however, the consistency of feature importance via different methods remains understudied . |
| Approach: | They compare feature importance from built-in mechanisms and post-hoc methods that approximate model behavior to find similarities between models. |
| Outcome: | The proposed methods show that features from traditional models are more similar with each other than with deep learning models. |
Decision-Focused Summarization (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing summarization methods define relevance based on textual information alone without incorporating insights about a particular decision. |
| Approach: | They propose a method that summarizes relevant information for a decision using full text . they then build a model that makes the decision based on the full text while accounting for textual non-redundancy. |
| Outcome: | The proposed method outperforms text-only summarization methods and model-based explanation methods in decision faithfulness and representativeness. |
Characterizing the Value of Information in Medical Notes (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Obtaining and analyzing information is critical for the diagnosis, prognosis, treatment, and prevention of disease. |
| Approach: | They propose a probing framework to select parts of notes that enable more accurate predictions than using all notes. |
| Outcome: | The proposed framework achieves better predictive performance with only 6.8% of all tokens for readmission prediction. |
Ecologically Valid Explanations for Label Variation in NLI (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Human label variation exists in many natural language processing tasks, including NLI . |
| Approach: | They build an English dataset of 1,415 ecologically valid explanations for 122 MNLI items . they find that people can systematically vary on their interpretation . |
| Outcome: | The proposed dataset contains 1,415 ecologically valid explanations for 122 items . the results show that people can vary on interpretation and highlight differences . |
No Permanent Friends or Enemies: Tracking Relationships between Nations from News (N19-1)
Copied to clipboard
| Challenge: | Understanding complex international relations is important but challenging for civilians . topic models and neural models have been proposed to explore relations without supervision . |
| Approach: | They propose an unsupervised neural model that integrates linguistic insights into the model to infer relations between nations from news articles. |
| Outcome: | The proposed model outperforms baselines from topic models and hidden Markov models. |