Challenge: a failure in one tool can trigger a cascade of errors, leading to complete task failure.
Approach: They propose a framework for tools more broadly which explores a model’s ability to detect “silent” tool errors and reflect on how to plan.
Outcome: The proposed approach shows that the model can detect "silent" tool errors and plan.

Similar Papers

Benchmarking Failures in Tool-Augmented Language Models (2025.naacl-long)

Copied to clipboard

Challenge: FAIL-TaLMs contains 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage.
Approach: They introduce a benchmark to examine the shortcomings of tool-augmented language models (TaLMs) that assume 'perfect' information access and tool availability.
Outcome: The proposed benchmark systematically evaluates 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage.
Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: emergence of tool agent paradigm has broadened capability boundaries of the Large Language Model (LLM) but effectiveness of tool agents limited due to parameter failure during execution .
Approach: They propose a parameter failure taxonomy to investigate parameter failure . they propose suggestions for standardizing tool return formats and improving error feedback mechanisms .
Outcome: The proposed model is based on a tool agent invocation chain and a mainstream tool agent . it shows that parameter name hallucination failure stems from inherent limitations .
FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models are being increasingly deployed as decision-making core of autonomous agents . however, in conversational benchmarks, these agents fail due to the cascading effects of incorrect decision- making .
Approach: They propose a framework that analyzes failure trajectories from baseline agents to identify most prevalent errors.
Outcome: Experiments show that the framework improves performance over open-source LLMs . the framework can be used to build reliable, multi-turn tool-use agents .
ACEBench: A Comprehensive Evaluation of LLM Tool Usage (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for evaluating LLMs’ tool usage face several limitations: limited evaluation scenarios, lacking assessments in real multi-turn dialogue contexts; narrow evaluation dimensions, with insufficient detailed assessments of how LLM use tools; and reliance on LLM or real API executions for evaluation, which introduces significant overhead.
Approach: ACEBench is a benchmark for evaluating tool usage in Large Language Models . it categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
Outcome: ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
Tool Preferences in Agentic LLMs are Unreliable (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) can now access a wide range of external tools thanks to the Model Context Protocol (MCP).
Approach: They expose a vulnerability in prevalent tool/function-calling protocols by editing tool descriptions to find out which tools are used by LLMs.
Outcome: The proposed changes in the tool descriptions can increase the usage of tools from LLMs when competing with alternatives.
Learning to Ask: When LLM Agents Meet Unclear Instruction (2025.emnlp-main)

Copied to clipboard

Challenge: Despite their impressive capabilities, LLMs struggle with complex computations and delivering accurate, timely information.
Approach: They propose a framework that prompts LLM agents to ask questions when they encounter obstacles due to unclear instructions and an automated evaluation tool called ToolEvaluator.
Outcome: The proposed framework outperforms existing frameworks for tool learning in the Noisy ToolBench.
Thesis Proposal: When Does an Agent Know It Is Lost? Confidence Trajectory Analysis for Tool-Using LLMs (2026.acl-srw)

Copied to clipboard

Challenge: Existing uncertainty quantification methods treat each step in isolation, ignoring how confidence evolves and compounds across a full task trajectory.
Approach: They propose a framework for trajectory-level confidence analysis in the tool-use agent setting.
Outcome: The proposed framework will expose early warning signals for agent failure and offer interpretable diagnostic tools for understanding when and why LLM agents lose confidence.
ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering (2026.acl-long)

Copied to clipboard

Challenge: Large language model (LLM) agents often face strict input context limits, preventing efficient consideration of large toolsets.
Approach: They propose a tool that allows LLMs to merge tools with auto-correction and toolScopeRetriever to rank and select only the most relevant tools for each query.
Outcome: Evaluations on three state-of-the-art LLMs and three open-source tool-use benchmarks show gains of 8.38% to 38.6% in tool selection accuracy.
A Review of Prominent Paradigms for LLM-Based Agents: Tool Use, Planning (Including RAG), and Feedback Learning (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models have been used for planning, tool use, and feedback learning . inconsistent taxonomy and complexity of workflows create challenges .
Approach: They propose a unified taxonomy to review and discuss the three paradigms . they define environments/tasks, common LLM-profiled roles and universally applicable workflows based on prior work .
Outcome: The proposed taxonomy compares LMPR implementations and workflow usage across paradigms . large language models have human-like reasoning capabilities, the authors say .
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents (2026.acl-long)

Copied to clipboard

Challenge: a fundamental pillar of trustworthiness is calibration, which refers to an agent’s ability to express confidence that reliably reflects its actual performance.
Approach: They propose a reinforcement learning framework that jointly optimizes task accuracy and calibration, supported by a holistic benchmark of reward designs.
Outcome: The proposed framework improves calibration across tool types and shows that trained agents achieve superior calibration and exhibit robust generalization from local training environments to noisy web settings and to distinct domains such as mathematical reasoning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations