Papers with ELO

5 papers
Personalized Benchmarking: Evaluating LLMs by Individual Preferences (2026.findings-acl)

Copied to clipboard

Challenge: Current benchmarks average preferences across all users to compute aggregate ratings . this overlooks individual user preferences when establishing model rankings .
Approach: They compute personalized model rankings using ELO ratings and Bradley-Terry coefficients . they find users exhibit substantial heterogeneity in topical interests and communication styles .
Outcome: The results show that individual rankings of LLM models diverge dramatically from aggregate rankings . a compact combination of topic and style features provides a useful feature space .
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs (2026.eacl-industry)

Copied to clipboard

Challenge: Recent studies have focused on enhancing multilingual large language models (MLLMs) for specific languages.
Approach: They propose an efficient layer-specific optimization method to enhance continual pretraining (CP) for specific languages in multilingual large language models (MLLMs).
Outcome: The proposed method achieves a training speedup of up to 6.46 times compared to existing methods while improving target language performance by up to 5.2% on qualitative benchmarks.
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference (2025.findings-naacl)

Copied to clipboard

Challenge: Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences.
Approach: They propose a system-level evaluation framework that ranks LLMs based on their alignment with human preferences.
Outcome: The proposed framework aims to rank LLMs based on their performance and alignment with human preferences.
Compare without Despair: Reliable Preference Evaluation with Generation Separability (2024.findings-emnlp)

Copied to clipboard

Challenge: a meta-evaluation measure, separability, estimates how suitable a test instance is for pairwise preference evaluation.
Approach: They propose a measure of separability which measures how suitable a test instance is for pairwise preference evaluation.
Outcome: The proposed measure shows that instances with high separability yield more consistent preference ratings from human- and auto-raters.
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat (2025.acl-long)

Copied to clipboard

Challenge: Evaluating large language models (LLMs) is a complex task. Pairwise ranking has emerged as state-of-the-art method to evaluate human preferences.
Approach: They propose to use pairwise ranking to evaluate human preferences . they propose to evaluate the robustness of ranking algorithms in LLMs .
Outcome: The proposed methods are based on the principles of effective ranking and the robustness of several ranking algorithms in the context of LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations