Papers by Chao Wen
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to predict human emotions are inconsistent due to complexity of emotion and subjectivity of perception. |
| Approach: | They propose a Bayesian approach to estimate uncertainty in emotion attributes using a deep neural network model. |
| Outcome: | The proposed approach estimates uncertainty in emotion attributes along with aleatoric and epistemic uncertainties. |
Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution Estimation (2024.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that incorporating utterances without majority-agreed labels into an additional class reduces the classification performance of the other emotion classes. |
| Approach: | They propose to combine utterances without majority-agreed labels into an additional class . they propose to quantify uncertainty in emotion classification using evidential deep learning . |
| Outcome: | The proposed method retains classification accuracy while effectively detects ambiguous emotion expressions. |
Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing watermarking methods embed imperceptible identifiers into text to address copyright concerns. |
| Approach: | They propose a new philosophy for watermark attacks that addresses watermark collision . they demonstrate that collision poses a threat to all logit-based watermark algorithms . |
| Outcome: | The proposed method improves watermark collision performance on top of other methods. |
Attribution-Based Analysis and Optimization of Modular Agentic Workflows (2026.findings-acl)
Copied to clipboard
Yingxuan Yang, Bo Huang, Siyuan Qi, Chao Feng, Haoyi Hu, Yuxuan Zhu, Jinbo Hu, Haoran Zhao, Ziyi He, Xiao Liu, ZongYu Wang, Muning Wen, Lin Qiu, Xuezhi Cao, Xunliang Cai, Yong Yu, Weinan Zhang
| Challenge: | Large Language Models (LLMs) have driven the rise of agentic workflows . yet, how can we attribute performance gains to individual upgrades and their interactions? |
| Approach: | They propose a game-theoretic framework that models component upgrades as players and evaluates component coalitions to compute Shapley values. |
| Outcome: | The proposed framework provides interaction-aware attribution and recommendation for model allocation under a fixed workflow structure. |
Training Language Model to Critique for Better Refinement (2025.findings-acl)
Copied to clipboard
Tianshu Yu, Chao Xiang, Mingchuan Yang, Pei Ke, Bosi Wen, Cunxiang Wang, Jiale Cheng, Li Zhang, Xinyu Mu, Chuxiong Sun, Minlie Huang
| Challenge: | Large language models (LLMs) have remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. |
| Approach: | They propose a framework to train critic models using refinement signals to generate feedback loops where critiques guide the model in refining its responses. |
| Outcome: | The proposed framework outperforms traditional methods and open-source models in terms of critique quality and refinement outcomes. |
Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment (2025.acl-long)
Copied to clipboard
| Challenge: | Large language and multimodal models have shown remarkable success on various benchmarks focused on specific skills such as general-purpose programming, math word problem-solving, and visual question answering. |
| Approach: | They propose a program synthesis benchmark based on real-world programming tasks . they propose 'fine-tuning pipeline' to boost performance of large language models . |
| Outcome: | The proposed model outperforms existing models on tasks that require a combination of skills on visual programming and programming. |
Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom’s Taxonomy (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods suffer from cognitive dimensional simplification and methodological unreliability due to the ”LLM-as-a-Judge” approach. |
| Approach: | They propose a six-tiered benchmark that evaluates ASG systems by prioritizing deterministic algorithms and introducing a GRADE approach for abstract abilities. |
| Outcome: | The proposed method provides the ASG field with a systematic, reproducible, and theoretically grounded benchmark to guide future research. |
Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for multi-role dialogue summarization favor surface-level imitation of references rather than genuine gains in faithfulness or alignment with human preferences. |
| Approach: | They propose a framework that couples explicit cognitive-style reasoning with reward-based optimization for multi-role dialogue summarization. |
| Outcome: | The proposed framework matches strong baselines on ROUGE and BERTScore, while in-depth analysis on SAMSum shows clear gains in factual faithfulness and model-based preference alignment. |
Modelling Variability in Human Annotator Simulation (2024.findings-acl)
Copied to clipboard
| Challenge: | Human annotator simulation (HAS) is a cost-effective alternative to human evaluation tasks. |
| Approach: | They propose a framework to model human annotation variability via meta-learning . conditional softmax flow model leverages diverse human annotations via meta learning . results demonstrate that method can predict aggregated behaviours of human annotators . |
| Outcome: | The proposed method achieves state-of-the-art performance on two real-world human evaluation tasks: emotion recognition and toxic speech detection. |
TURTLEAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics (2026.findings-acl)
Copied to clipboard
| Challenge: | Vision-language models have been explored for visual programming, but performance is unclear . most prior work focuses on visual programming for productivity . |
| Approach: | They propose a visual programming benchmark that uses visual programming to evaluate VLMs. |
| Outcome: | The proposed model improves on GPT-5, GPT-4o, and Qwen2-VL-72B on real-world tasks by 20% . the proposed model is based on 823 visual programming tasks in the Turtle Graphics domain . |