Papers by Chenjun Xu
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation (2026.acl-long)
Copied to clipboard
Katelyn X. Mei, Yi-Li Hsu, Minjoon Choi, Zongwan Cao, Chenjun Xu, Bingbing Wen, Su Lin Blodgett, Lucy Lu Wang
| Challenge: | a large-scale analysis of human evaluation protocols for long-form generation tasks is lacking in current practice . current protocols lack proper standardization and operationalization, which can limit validity of evaluation . |
| Approach: | They conduct a large-scale analysis of human evaluation protocols for long-form generation tasks in *CL conference papers from 2023–2025. |
| Outcome: | The proposed evaluation protocols lack standardization and operationalization, the authors show . they also find that the evaluation protocols are inadequate for specific domains and tasks . |
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Psychology research has shown that humans are poor at estimating their performance on tasks, tending towards underconfidence on easy tasks and overconfidence on difficult tasks. |
| Approach: | They propose to use a self-assessment method to assess confidence in large language models (LLMs) they propose to ask for the answer separately and then use them to improve their accuracy. |
| Outcome: | The proposed method improves confidence calibration and interpretability in QA tasks with different personas. |