Papers by Muyu He
TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games (2025.emnlp-main)
Copied to clipboard
| Challenge: | evaluating large language models' reasoning abilities via detective stories is often infeasible due to the large answer space and diverse reasoning types presented by its questions. |
| Approach: | They propose a framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. |
| Outcome: | The proposed framework and dataset are based on the detective games Ace Attorney and Danganronpa and show that they are more efficient than current strategies for enhancing deductive reasoning. |
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Small shifts in user behavior can cause sharp drops in agent performance . prior work has shown that LLMs lack robustness to real-world noise and small input perturbations. |
| Approach: | They propose a model-agnostic method for systematically stress testing AI agents that learns directions in activation space corresponding to steerable user traits. |
| Outcome: | The proposed method can be used to stress test AI agents in airline, retail, telecom, and telehealth domains. |