Challenge: Current research emphasizes LLMs’ capacity to utilize tools in well-structured environments while overlooking their stability when confronted with the inevitable noise of the real world.
Approach: They propose a multi-level benchmark to evaluate the robustness of large language models in tool learning by establishing five external environments with varying levels of noise.
Outcome: The proposed model outperforms the GPT-4 model in tool learning in three critical phases: tool selection, parameter identification, and content filling.

Similar Papers

StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning.
Approach: They propose a virtual API server and stable evaluation system to assess the stability of large-scale real-time APIs.
Outcome: The proposed benchmarks demonstrate the stability of the proposed system and its caching system.
ACEBench: A Comprehensive Evaluation of LLM Tool Usage (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks for evaluating LLMs’ tool usage face several limitations: limited evaluation scenarios, lacking assessments in real multi-turn dialogue contexts; narrow evaluation dimensions, with insufficient detailed assessments of how LLM use tools; and reliance on LLM or real API executions for evaluation, which introduces significant overhead.
Approach: ACEBench is a benchmark for evaluating tool usage in Large Language Models . it categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
Outcome: ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent.
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation datasets lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage.
Approach: They propose to use multilingual consistency as a complementary metric to assess performance bottlenecks and guide model improvement.
Outcome: The proposed model lacks cross-lingual alignment and language coverage gaps between state-of-the-art models.
LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing SEA-focused benchmarks miss Lao-specific cultural grounding and linguistic properties.
Approach: They propose a multi-dimensional benchmark for assessing large language models in Lao . they use open-source and held-out subsets to evaluate languages with a hybrid pipeline .
Outcome: LaoBench is the first large-scale, high-quality, and multidimensional benchmark for assessing LLM language understanding and reasoning in Lao.
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches fail to fully capture all risks in tool utilization, resulting in financial loss or privacy leaking.
Approach: They propose a framework to assess the safety of LLM tool utilization in a prospective manner, covering malicious user instructions and diverse practical toolsets.
Outcome: The proposed framework significantly enhances LLMs’ self-awareness, enabling a more safer and trustworthy tool utilization.
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations assume tool use in short contexts, offering limited insight into model behavior during realistic long-term interactions.
Approach: a benchmark is a tool to test long-term tool use in large language models . the tool includes multiple tasks execution contexts and realistic noise .
Outcome: a new benchmark tests the tool use capabilities in long-term interactions.
UniToolBench: A Benchmark for Tool-Augmented LLMs in Cross-Domain, Universal Task Automation (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks that focus on manually curated tool graphs lack scalability and diversity across domains.
Approach: They propose a large-scale, cross-domain benchmark to evaluate LLMs' ability to reason over and utilize interconnected tools for automation.
Outcome: The proposed benchmark incorporates automated tool graph construction by formulating link prediction as a probabilistic task, instead of relying on categorical LLM outputs.
SafeLawBench: Towards Safe Alignment of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies indicate that large language models (LLMs) may exhibit risks, including threats to the protection of private data and the generation of hallucinations.
Approach: They propose to evaluate LLMs from a legal perspective using the SafeLawBench benchmark.
Outcome: The proposed framework categorizes safety risks into three levels based on legal standards and includes 24,860 multi-choice questions and 1,106 open-domain question-answering tasks.
VoiceBench: Benchmarking LLM-Based Voice Assistants (2026.tacl-1)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have enabled real-time speech interactions through LLMs.
Approach: They propose a benchmark specifically designed to assess LLM-based voice assistants.
Outcome: The proposed benchmark measures the performance of LLM-based voice assistants across eight tasks.
TRUEBench: Can LLM Response Meet Real-world Constraints as Productivity Assistant? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks fail to evaluate large language models' instruction-following capabilities . current benchmarks lack multilinguality, implicit constraints and multi-turn dialogue .
Approach: a new benchmark is designed to evaluate large language models' instruction-following capabilities . the benchmark features input prompts across 12 languages and includes inter-instance multilingual instructions .
Outcome: a new benchmark for large language models (LLMs) is designed to assess their performance in real-world settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations