Papers with WildBench
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal (2025.emnlp-main)
Copied to clipboard
| Challenge: | a recent study shows that asking for direct user feedback can be disruptive . we examine whether incorporating the contents of user feedback improves model performance . |
| Approach: | They analyze user feedback in the user-LLM conversation logs and harvest learning signals from it. |
| Outcome: | The proposed approach can lead to model degradation on two user-LM interaction datasets. |
IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to optimize conversational agents often rely on explicit preference pairs and expert evaluations. |
| Approach: | They propose a conversational agent framework that leverages the structured dependency between agent responses and user reactions to extract implicit feedback. |
| Outcome: | The proposed framework improves on MT-Bench-101, WildBench, and FB-Bech, and shows that mining implicit feedback supports better multi-turn alignment under evolving user preferences. |
TOWER: Tree Organized Weighting for Evaluating Complex Instructions (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Evaluating the ability of large language models to follow human-written instructions remains a challenge. |
| Approach: | They propose a new evaluation metric that incorporates human-judged importance into the assessment of complex instruction following. |
| Outcome: | The proposed evaluation metric incorporates human-judged importance into the assessment of complex instruction following. |