Papers by Jeongbeom Jeong
Arena-lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current benchmarks typically compare system outputs against baselines, but this method yields lower reliability than direct comparison. |
| Approach: | They propose to integrate tournament structure on top of head-to-head comparison. |
| Outcome: | The proposed model achieves higher reliability with fewer comparisons even with smaller datasets or weaker judges. |