Papers by Georgii Levtsov
Confidence and Stability of Global and Pairwise Scores in NLP Evaluation (2025.acl-srw)
Copied to clipboard
| Challenge: | Modern natural language processing benchmarks are often represented as pairwise comparison leaderboards, such as LMSYS Arena. |
| Approach: | They investigate the strengths and weaknesses of global scores and pairwise comparisons to aid decision-making in selecting appropriate model evaluation strategies. |
| Outcome: | The proposed method underestimates strong models with rare errors or low confidence, while relying on global scores can be more effective. |