Papers by Yuhan Wang
Make Your Decision Convincing! A Unified Two-Stage Framework: Self-Attribution and Decision-Making (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing frameworks for explaining black-box model behavior are unreliable . large-scale pre-trained models often rely on superficial clues for predictions . |
| Approach: | They propose a unified two-stage framework that uses subsequences from the input text as a rationale to generate model decision. |
| Outcome: | The proposed framework achieves competitive results on five reasoning datasets and in semi-supervised scenarios. |
Retrieval Models Aren’t Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) suffer from inherent inabilities to interact with the physical world and access vast, up-to-date knowledge. |
| Approach: | They propose a tool retrieval benchmark for large language models (LLMs) that includes 7.6k diverse retrieval tasks and a corpus of 43k tools. |
| Outcome: | The proposed model performs poorly on the heterogeneous tool retrieval benchmark, resulting in low pass rate and low retrieval quality. |
CPT-Agent: A Cognitive Process Theory-driven Framework for Student Simulation in Writing Development (2026.acl-long)
Copied to clipboard
| Challenge: | Existing LLMs model overly capable learners who over-apply feedback, resulting in pedagogically implausible behavior. |
| Approach: | They propose a framework that decouples cognitive ability from writing proficiency and models their interaction during writing and revision. |
| Outcome: | The proposed model produces distinguishable proficiency levels and is consistent with instructional theories. |
AS-ES Learning: Towards efficient CoT learning in small models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to induce Chain-of-Thought (CoT) in LLMs are limited and do not consider the importance of efficiently utilizing existing CoT data. |
| Approach: | They propose a new training paradigm which exploits the inherent information in CoT for iterative generation. |
| Outcome: | The proposed training paradigm surpasses direct seq2seq training on CoT-extensive tasks without data augmentation or altering the model itself. |
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness (2026.acl-long)
Copied to clipboard
Jingyu Lu, Yuhan Wang, Fan Zhuo, Xize Cheng, Changhao Pan, Xueyi Pu, Yifu Chen, Chenyuhao Wen, Tianle Liang, Zhou Zhao
| Challenge: | SDiaReward is an end-to-end spoken dialogue system that integrates paralinguistic nuances and spontaneous nature of human conversation. |
| Approach: | They propose an end-to-end multi-turn reward model trained on SDiaReward-Dataset . it is a collection of episode-level preference pairs targeting modality and colloquiality gaps . |
| Outcome: | The proposed model outperforms general-purpose audio LLMs in episode-level evaluation. |
Is He Extroverted? Identifying Missing Relevant Personas for Faithful User Simulation (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing user simulation approaches focus on generating user-like responses in dialogue without verifying whether critical personas are supplied. |
| Approach: | They propose a task of identifying persona dimensions that are relevant but missing in simulating a user's reply for a given dialogue context. |
| Outcome: | The proposed model identifies persona dimensions that are relevant but missing in simulating a user’s response for a given dialogue context. |
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning (2026.acl-long)
Copied to clipboard
| Challenge: | Recent latent reasoning methods suffer from a bandwidth bottleneck . explicit textual rationales suffer from premature semantic collapse . |
| Approach: | They propose a new paradigm that reformulates visual deduction via Dynamic Windowed Alignment Learning. |
| Outcome: | The proposed paradigm achieves state-of-the-art performance among latent reasoning methods surpassing the strong baseline Monet by 5.03% on average. |
STORM-BORN: A Challenging Mathematical Derivations Dataset Curated via a Human-in-the-Loop Multi-Agent Framework (2025.findings-acl)
Copied to clipboard
Wenhao Liu, Zhenyi Lu, Xinyu Hu, Jerry Zhang, Dailin Li, Jiacheng Cen, Huilin Cao, Haiteng Wang, Yuhan Li, Xie Kun, Dandan Li, Pei Zhang, Chengbo Zhang, Yuxiang Ren, Xiaohong Huang, Yan Ma
| Challenge: | Existing datasets suffer from outdated and insufficient challenging content, neglecting human-like reasoning, and limited reliability due to single-LLM generation. |
| Approach: | They propose a human-in-the-loop, multi-agent data generation framework that integrates reasoning-dense filters, multiagent collaboration, and human mathematicians’ evaluations to ensure the reliability and quality of the dataset. |
| Outcome: | The proposed framework improves accuracy and quality of the 2,000-synthesized datasets by integrating reasoning-dense filters, multi-agent collaboration, and human mathematicians’ evaluations. |
Improving Empathetic Response Generation by Recognizing Emotion Cause in Conversations (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to empathetic response generation ignore the emotion cause . existing dialogue systems lack emotion understanding and empathy . |
| Approach: | They propose a framework that integrates emotion cause information into empathetic response generation by predicting context emotion labels and sequence of emotion cause-oriented labels. |
| Outcome: | The proposed framework improves empathetic response generation by incorporating emotion cause information into the model. |
DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedback (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Multi-agent systems (MAS) powered by large language models struggle to adapt to evolving task dependencies and to handle uncertainties. |
| Approach: | They propose a Dynamic Environment-Aware Manager-Player Agents Coordination framework that enhances multi-agent coordination through long-term strategic planning. |
| Outcome: | The proposed framework outperforms traditional reinforcement learning and human-agent collaboration in the Overcooked simulation. |
Attention Weights as an Indicator: Analyzing and Improving Document Utilization in Retrieval-Augmented Generation (2026.acl-long)
Copied to clipboard
| Challenge: | In traditional RAG models, documents are grouped into categories based on their quality and order, and the quality of inputs is variable due to ineffective retrievers or misalignment between the retriever and generator. |
| Approach: | They propose to use attention weights to enhance document utilization from three perspectives: document ranking, placement, and filtering. |
| Outcome: | The proposed method outperforms baselines and improves document utilization effectiveness in a training-free manner. |
Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs (2026.findings-acl)
Copied to clipboard
Linhao Zhang, Yuhan Song, Aiwei Liu, Chuhan Wu, Sijun Zhang, Wei Jia, Yuan Liu, Houfeng Wang, Zhou Xiao
| Challenge: | Recent Audio Large Language Models (AudioLLMs) excel at reasoning tasks, but struggle at elementary auditory perception. |
| Approach: | They propose a framework that organizes audio information into three explicit components in a unified JSON format. |
| Outcome: | The proposed framework boosts fine-grained perception by 10.9% on MMSU over state-of-the-art models while preserving robust reasoning capabilities. |
Act2P: LLM-Driven Online Dialogue Act Classification for Power Analysis (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies focus on explicit utterance functions, overlooking the implicit power dynamics embedded in dialogue. |
| Approach: | They propose an online Dialogue Act Classification and Dynamic Power Analysis framework based on large language models to integrate dialogue act classification with power quantification. |
| Outcome: | The proposed framework outperforms existing methods in online scenarios and shows that dialogue power is distributed and dynamically transferred. |
Would You Like to Make a Donation? A Dialogue System to Persuade You to Donate (2024.lrec-main)
Copied to clipboard
| Challenge: | Persuasive automated dialogue systems are a popular way to influence people's behavior and decision making. |
| Approach: | They propose to use a context-aware persuasion strategy selection module to persult users . they also propose a persuasiveness prediction model to automatically evaluate the persuasiveness of generated text. |
| Outcome: | The proposed system can achieve better performance on several automated evaluation metrics than baseline models. |
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis (2025.findings-acl)
Copied to clipboard
Yu Zhang, Wenxiang Guo, Changhao Pan, Dongyu Yao, Zhiyuan Zhu, Ziyue Jiang, Yuhan Wang, Tao Jin, Zhou Zhao
| Challenge: | Existing zero-shot singing voice synthesis models depend on phoneme and note boundary annotations, limiting their robustness and producing poor transitions between phonemes and notes. |
| Approach: | They propose a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. |
| Outcome: | Experimental results show that TCSinger 2 outperforms baseline models in subjective and objective metrics across multiple related tasks. |
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety (2025.emnlp-main)
Copied to clipboard
Jiahao Qiu, Yinghui He, Xinzhe Juan, Yimin Wang, Yuhan Liu, Zixin Yao, Yue Wu, Xun Jiang, Ling Yang, Mengdi Wang
| Challenge: | EmoAgent evaluates and mitigates mental health hazards in human-AI interactions, especially for vulnerable human users with psychological disorders. |
| Approach: | EmoAgent is a multi-agent AI framework designed to evaluate and mitigate mental health hazards in human-AI interactions. |
| Outcome: | EmoAgent evaluates and mitigates mental health hazards in human-AI interactions. |
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have catalyzed numerous AI applications, among which role-playing agents (RPAs) are particularly popular. |
| Approach: | They propose to evaluate LLMs' character understanding capability via the character profiling task, i.e., summarizing character profiles from corresponding materials, a widely adopted yet understudied practice for RPA development. |
| Outcome: | The proposed model outperforms existing models and literature summarization methods and proves its ability to understand fictional characters in downstream tasks. |
H-MAS: Hierarchical Multi-Agent Scheduling for Multi-Tenant LLM Serving (2026.findings-acl)
Copied to clipboard
Yuhan Liu, Cong Xu, Qi Jia, Yihua Wang, Feiyu Chen, Liang Jin, Lu Liu, Yaqian Zhao, Yuting Ding, Xiang Li
| Challenge: | Multi-tenant Model-as-a-Service (MaaS) workloads exhibit non-stationarity across multiple time scales . existing request schedulers often rely on a fixed policy that remains unchanged at runtime . |
| Approach: | They propose a hierarchical multi-agent scheduler that operates in a layered closed loop . they propose to maintain 1.2–3.0 higher Goodput than SGLang and vLLM . |
| Outcome: | Experiments show that H-MAS achieves 1.2–3.0 higher Goodput than SGLang and vLLM . it maintains more stable QoS under diverse request lengths and heterogeneous SLO targets . |
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing positional encodings exhibit long-term decay, based on an entrenched and long-standing opinion that tokens farther away from the current position carry less relevant information. |
| Approach: | They propose a high-frequency rotary position encoding (HoPE) that replaces specific components in RoPE with position-independent ones, retaining only high- frequency signals. |
| Outcome: | The proposed method exhibits greater robustness to the out-of-distribution behavior in attention patterns during extrapolation. |
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models (2024.findings-acl)
Copied to clipboard
Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, Junran Peng
| Challenge: | Large Language Models (LLMs) have paved the way for complex tasks such as role-playing. |
| Approach: | They propose a framework to benchmark, elicit, and enhance role-playing abilities in Large Language Models. |
| Outcome: | The proposed framework improves role-playing abilities with 168,093 samples. |
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning (2026.acl-long)
Copied to clipboard
Bowen Ding, Yuhan Chen, Jiayang Lyu, Jiyao Yuan, Qi Zhu, Shuangshuang Tian, Dantong Zhu, Futing Wang, Heyuan Deng, Fei Mi, Lifeng Shang, Tao Lin
| Challenge: | evolving generic Large Language Models into specialized Large Reasoning Models requires effective post-training. |
| Approach: | They propose a plasticity-ceiling framework to harness expert trajectories . they establish the Sequential SFT-then-RL pipeline as the superior standard . |
| Outcome: | The proposed framework overcomes stability and premature convergence deficits in synchronized approaches. |
AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models are increasingly used to automate data analysis, but data science tasks often admit multiple statistically valid solutions. |
| Approach: | They propose a framework to evaluate LLM-generated code and assess its reproducibility . they introduce two reproducibility-enhancing prompting strategies and benchmark them against standard prompting . |
| Outcome: | The proposed framework improves reproducibility of large language models . it provides a foundation for transparent, reliable, and efficient human–AI collaboration in data science. |
GlossaGen: Making Academic Translation Smarter with Glossing (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing machine translation systems obscure or mistranslate key terminology, while paraphrasing aimed at lay readers often oversimplifies it, hindering their ability to master domain-specific technical vocabulary. |
| Approach: | They propose a task which produces translations dynamically adapted to a reader’s academic proficiency, or level, and a framework to address this challenge. |
| Outcome: | The proposed framework achieves higher scores than baselines on a synthesized benchmark and human evaluations. |
Large Language Models Meet Harry Potter: A Dataset for Aligning Dialogue Agents with Characters (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models that can create open-domain dialogue agents lack character representation and annotations. |
| Approach: | They propose a dataset to study character alignment and character representation . it includes all dialogue sessions from the Harry Potter series and includes annotations . |
| Outcome: | The proposed dataset can be used as a universal benchmark for character-driven LLMs. |