Papers by Zheyuan Zhang

23 papers
Do LLMs Catch Their Own Mistakes? A Comprehensive Benchmark for Reflective Tool Use LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks primarily evaluate planning and execution success, overlooking the self-reflective dimension of tool use.
Approach: They propose a benchmark to assess LLMs’ self-reflective reasoning in tool-augmented multi-turn dialogues.
Outcome: The proposed benchmark covers 10 domains with 88 distinct APIs and 968 annotated dialogues, systematically injecting diverse error types arising from both user and assistant behavior.
Interpretable Graph-Language Modeling for Detecting Youth Illicit Drug Use (2026.findings-eacl)

Copied to clipboard

Challenge: Illicit drug use among teens and young adults remains a public health concern . existing models ignore latent and interconnected structures among survey variables .
Approach: They propose a joint graph-language modeling framework to detect illicit drug use among TYAs . they use large-scale surveys such as the Youth Risk Behavior Survey and the National Survey on Drug Use and Health to analyze data .
Outcome: The proposed framework outperforms baseline models on YRBS and NSDUH datasets in predictive accuracy.
From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive performance in various language tasks, but are prone to spurious correlations and illusory information.
Approach: They propose to use pre-trained language models to justify decisions with formalized, coherent reasoning chains.
Outcome: The proposed strategies improve coherence of rationalizations yielding state-of-the-art results on Tiered Reasoning for Intuitive Physics (TRIP).
Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment Approach (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on LLMs evaluation with exams are lacking in cognitive research on their overall knowledge structure.
Approach: They conduct an evaluation using a human test dataset based on Bloom Taxonomy to reveal the knowledge structures of Large Language Models and gain insights of their cognitive capabilities.
Outcome: The proposed model can pass AP, SAT, and Leetcode exams, but lacks the cognitive power to perform on human exams.
Instant Personalized Large Language Model Adaptation via Hypernetwork (2026.acl-long)

Copied to clipboard

Challenge: Existing parameter-efficient fine-tuning methods require training a separate adapter for each user, making them computationally expensive and impractical for real-time updates.
Approach: They propose a scalable framework that maps a user's profile directly to a full set of adapter parameters.
Outcome: The proposed framework outperforms prompt-based personalization and OPPU while using substantially fewer computational resources at deployment.
AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to agent routing emphasize cost efficiency while overlooking the fine-grained contextual and relational structure inherent in QA tasks.
Approach: They propose a framework that formulates multi-agent QA as a knowledge-graph-guided routing problem supervised by empirical performance signals.
Outcome: The proposed framework outperforms single-agent and ensemble baselines while generalizing across benchmarks and LLM backbones.
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) have advanced from perception tasks to complex multi-step reasoning.
Approach: They propose a framework that integrates reinforcement learning with verifiable rewards with process-level supervision through automatically collected rubric-based generative rewards.
Outcome: The proposed framework achieves state-of-the-art performance on six multimodal reasoning benchmarks and significantly improves reasoning faithfulness in dedicated evaluations.
NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Diet plays a critical role in human health, but tailoring dietary reasoning to individual health conditions remains a challenge.
Approach: a new benchmark evaluates dietary reasoning using a national health survey data set.
Outcome: The NGQA benchmark evaluates dietary reasoning across three tasks using a set of question complexity settings and baseline models.
A Combinatorial Approach to Neural Emergent Communication (2025.coling-main)

Copied to clipboard

Challenge: Existing research on emergent communication uses the Lewis signaling game . however, the training data is limited and the messages are often ineffective .
Approach: They propose a combinatorial algorithm to solve the symbolic complexity for classification, which is the minimum number of symbols in the message for successful communication.
Outcome: The proposed algorithm increases the number of effective symbols in the emergent language.
Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making (2025.findings-emnlp)

Copied to clipboard

Challenge: Modern embodied AI uses multimodal large language models as policy models, predicting actions from final-layer hidden states.
Approach: They propose a hierarchical action probing method that aggregates representations from all layers, mirroring the brain's multi-level organization.
Outcome: Experiments show that hierarchical probing improves on last-layer embodied models and achieves a 46.6% success rate and a 62.5% gain in spatial reasoning tasks.
LLM-Empowered Class Imbalanced Graph Prompt Learning for Online Drug Trafficking Detection (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to combat illicit drug trafficking are impractical due to the scarcity of labeled samples and imbalance of classes.
Approach: They propose a Large Language Model-empowered Heterogeneous Graph Prompt Learning framework for illicit drug trafficking detection that leverages LLM to facilitate heterogeneous graph neural networks to effectively identify minority classes.
Outcome: The proposed framework is able to identify minority classes in class-imbalanced scenarios.
Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models and Multimodal Large Language Modells can memorize sensitive information, raising ethical and privacy concerns.
Approach: They propose a novel unlearning framework that selectively clips neurons based on their relative importance to the targeted forget data.
Outcome: The proposed framework selectively clips neurons based on their relative importance to the targeted forget data, curated for different modalities.
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties (2024.emnlp-main)

Copied to clipboard

Challenge: Emergent In-context Learning on Videos induces in-contact learning over video and text . eILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.
Approach: They implement Emergent In-context Learning on Videos (EILeV) that induces in-contact learning over video and text by capturing key properties of pre-training data.
Outcome: The proposed training paradigm outperforms off-the-shelf VLMs in few-shot video narration for novel, rare actions.
MAPRO: Recasting Multi-Agent Prompt Optimization as Maximum a Posteriori Inference (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks.
Approach: They propose a framework that optimizes MAS prompts as a maximum a posteriori problem and then iteratively updates agent prompts.
Outcome: The proposed framework surpasses manual and automated benchmarks in multiple tasks and provides general guidelines for building more reliable and principled multi-agent systems in the future.
Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has attempted to mitigate this issue by using adaptive reasoning strategies, but these methods overlook a fundamental bottleneck: visual perception failures.
Approach: They propose a meta-reasoning controller that dynamically routes computation among three decision paths at each generation step.
Outcome: The proposed method outperforms slow-thinking methods while producing shorter responses.
Superficial Self-Improved Reasoners Benefit from Model Merging (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) rely heavily on large-scale reasoning data, but as data becomes scarce, model self-improvement offers a promising alternative.
Approach: They propose to merge the weights of original and self-improved LLMs to mitigate model collapse and improve generalized reasoning capability.
Outcome: The proposed model merge mitigates model collapse and improves generalized reasoning capability.
EmoBench: Evaluating the Emotional Intelligence of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Emotional Intelligence (EI) focus on emotion recognition, neglecting essential EI capabilities.
Approach: They propose a benchmark that proposes a comprehensive definition for machine EI . they propose 400 hand-crafted questions in English and Chinese to evaluate EI.
Outcome: The proposed benchmarks focus on emotion recognition, neglecting EI capabilities . they are constructed from existing datasets, which include frequent patterns and errors . the proposed benchmark includes questions in English and Chinese that require thorough reasoning and understanding .
NG-Router: Graph-Supervised Multi-Agent Collaboration for Nutrition Question Answering (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods for nutrition question answering face limited reasoning capacity and contextual overload . poor dietary patterns are associated with more than 11 million deaths in 2017 .
Approach: They propose a framework that enables supervised multi-agent collaboration for nutritional QA.
Outcome: The proposed framework outperforms single-agent and ensemble baselines in multi-agency reasoning tasks.
Can LLMs Convert Graphs to Text-Attributed Graphs? (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to model graph-structured data are limited by the availability of text-attributed graph data.
Approach: They propose a method to convert existing graphs into text-attributed graphs using large language models.
Outcome: The proposed method outperforms existing approaches that manually design node features on text-free graphs.
Simulating Classroom Education with LLM-Empowered Agents (2025.naacl-long)

Copied to clipboard

Challenge: Initial studies have focused on task-specific, independent LLM-empowered agents, but the potential of LLMs within a multi-agent collaborative framework for classroom simulation with real user participation remains unexplored.
Approach: They propose a multi-agent classroom simulation teaching framework that recognizes representative class roles and introduces a novel class control mechanism for automatic classroom teaching.
Outcome: The proposed framework can simulate dynamic learning environment for users with active teacher-student and student-studente interactions.
Behavior Knowledge Merge in Reinforced Agentic Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for supervised fine-tuning (SFT) are suboptimal to preserve task-specific capabilities on RL-trained agentic models.
Approach: They propose a distribution-aware merging framework specifically designed for RL-trained agentic models that disentangles shared and task-specific unique parameter updates while selectively preserving and rescaling unique ones.
Outcome: Experiments across multiple agent domains and model architectures show that the proposed framework surpasses baselines and unlocks synergistic potential among agents.
Transparent and Coherent Procedural Mistake Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Procedural mistake detection (PMD) is a problem of classifying whether a human user has successfully executed a task.
Approach: They extend PMD to require generating visual self-dialog rationales to inform decisions . they leverage a natural language inference model to formulate two automated metrics for coherence of generated rationale.
Outcome: The proposed model improves on a reframed task with a natural language inference model and a multi-faceted metrics visualization of common outcomes.
Explaining Length Bias in LLM-Based Preference Evaluations (2025.findings-emnlp)

Copied to clipboard

Challenge: a preference evaluation metric is often biased towards longer responses, revealing a reliability problem . a decomposition of the preference evaluation into two components is needed to understand this bias.
Approach: They propose to decompose the preference evaluation metric into two key components . the first component is length-dependent and related to trustworthiness .
Outcome: The proposed evaluation metric is based on two components: desirability and information mass.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations