Papers by Timon Willi
Balanced Accuracy: The Right Metric for Evaluating LLM Judges - Explained through Youden’s J statistic (2026.eacl-industry)
Copied to clipboard
| Challenge: | False refusals and task pass rates are key to reliable evaluation of large language models. |
| Approach: | They propose a principled best practice for evaluating judges based on a golden set of judge-quality metrics. |
| Outcome: | The proposed method improves the quality of judge-quality metrics on a golden set. |