Papers by James Ngai
Benchmarking Failures in Tool-Augmented Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | FAIL-TaLMs contains 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage. |
| Approach: | They introduce a benchmark to examine the shortcomings of tool-augmented language models (TaLMs) that assume 'perfect' information access and tool availability. |
| Outcome: | The proposed benchmark systematically evaluates 1,749 examples using 906 tools across 21 categories, including single- and multi-tool usage. |