Papers by Ben Slater
PredictaBoard: Benchmarking LLM Score Predictability (2025.findings-acl)
Copied to clipboard
Lorenzo Pacchiardi, Konstantinos Voudouris, Ben Slater, Fernando Martínez-Plumed, Jose Hernandez-Orallo, Lexin Zhou, Wout Schellaert
| Challenge: | Large Language Models (LLMs) fail unpredictably, demonstrating inconsistent success in even basic common sense reasoning tasks. |
| Approach: | They propose a framework to evaluate the ability of score predictors to anticipate LLM errors on specific task instances from existing datasets. |
| Outcome: | The proposed framework evaluates the ability of score predictors to anticipate LLM errors on specific task instances from existing datasets. |