Papers by Timon Willi

1 papers
Balanced Accuracy: The Right Metric for Evaluating LLM Judges - Explained through Youden’s J statistic (2026.eacl-industry)

Copied to clipboard

Challenge: False refusals and task pass rates are key to reliable evaluation of large language models.
Approach: They propose a principled best practice for evaluating judges based on a golden set of judge-quality metrics.
Outcome: The proposed method improves the quality of judge-quality metrics on a golden set.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations