Papers by Muyu He

2 papers
TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games (2025.emnlp-main)

Copied to clipboard

Challenge: evaluating large language models' reasoning abilities via detective stories is often infeasible due to the large answer space and diverse reasoning types presented by its questions.
Approach: They propose a framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa.
Outcome: The proposed framework and dataset are based on the detective games Ace Attorney and Danganronpa and show that they are more efficient than current strategies for enhancing deductive reasoning.
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents (2026.acl-long)

Copied to clipboard

Challenge: Small shifts in user behavior can cause sharp drops in agent performance . prior work has shown that LLMs lack robustness to real-world noise and small input perturbations.
Approach: They propose a model-agnostic method for systematically stress testing AI agents that learns directions in activation space corresponding to steerable user traits.
Outcome: The proposed method can be used to stress test AI agents in airline, retail, telecom, and telehealth domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations