Better than Average: Paired Evaluation of NLP systems (2021.acl-long)

Copied to clipboard

Challenge: Evaluation in NLP is usually done by comparing the scores of competing systems . averaging scores independently and declaring the best system is difficult .
Approach: They examine the use of averages to aggregate evaluation scores into a final number . they argue that the average ignores the pairing arising from the fact that systems are evaluated on the same test instances.
Outcome: The proposed method ignores the pairing arising from the fact that systems are evaluated on the same test instances.

Similar Papers

Confidence and Stability of Global and Pairwise Scores in NLP Evaluation (2025.acl-srw)

Copied to clipboard

Challenge: Modern natural language processing benchmarks are often represented as pairwise comparison leaderboards, such as LMSYS Arena.
Approach: They investigate the strengths and weaknesses of global scores and pairwise comparisons to aid decision-making in selecting appropriate model evaluation strategies.
Outcome: The proposed method underestimates strong models with rare errors or low confidence, while relying on global scores can be more effective.
Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status (2022.lrec-1)

Copied to clipboard

Challenge: comparing NLP systems by performance has become an essential question . comparing systems by performing performance criterion is criticized for allowing chance to determine superiority .
Approach: They propose to use bootstrap confidence intervals instead of state-of-the-art status and statistical significance testing to compare NLP system performance.
Outcome: The bootstrap confidence intervals are used to compare NLP system performance . the bootstrap test is more accurate than state-of-the-art status and statistical significance testing .
Active Evaluation: Efficient NLG Evaluation with Few Pairwise Comparisons (2022.acl-long)

Copied to clipboard

Challenge: Recent studies show that evaluating NLG systems using pairwise comparisons is expensive as the number of human annotations grows linearly with k.
Approach: They propose a framework to efficiently identify the top-ranked system by actively choosing system pairs for comparison using dueling bandit algorithms.
Outcome: The proposed framework reduces human annotations by 80% on 13 NLG evaluation datasets spanning 5 tasks .
What Can We Do to Improve Peer Review in NLP? (2020.findings-emnlp)

Copied to clipboard

Challenge: Traditionally, peer review is expected to act as a filter for high-quality, impactful work, but this does not hold in practice.
Approach: They argue that peer review is becoming increasingly spurious and that it is a problem for NLP . they propose a reproducibility checklist at EMNLP 2020 that could be used to ensure that papers are reproducible.
Outcome: The reproducibility checklist at EMNLP 2020 is the first step in that direction.
Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess Hypotheses (2020.acl-main)

Copied to clipboard

Challenge: Empirical research in natural language processing has adopted a narrow set of principles for assessing hypotheses . alternative approaches to assess hypothese rely on p-value computation, which suffers from several known issues.
Approach: They propose to compare different methods for assessing hypotheses . they argue that practitioners should first decide their target hypothesis before choosing a method .
Outcome: The proposed method differs from other methods, but is not widely used in NLP . the proposed method is based on a p-value computation, but has a small gap in accuracy .
Ranking Unraveled: Recipes for LLM Rankings in Head-to-Head AI Combat (2025.acl-long)

Copied to clipboard

Challenge: Evaluating large language models (LLMs) is a complex task. Pairwise ranking has emerged as state-of-the-art method to evaluate human preferences.
Approach: They propose to use pairwise ranking to evaluate human preferences . they propose to evaluate the robustness of ranking algorithms in LLMs .
Outcome: The proposed methods are based on the principles of effective ranking and the robustness of several ranking algorithms in the context of LLMs.
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference (2025.findings-naacl)

Copied to clipboard

Challenge: Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences.
Approach: They propose a system-level evaluation framework that ranks LLMs based on their alignment with human preferences.
Outcome: The proposed framework aims to rank LLMs based on their performance and alignment with human preferences.
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets (2021.acl-long)

Copied to clipboard

Challenge: Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measurements of harms.
Approach: They apply a measurement modeling lens to inventory pitfalls that threaten benchmarks' validity as measurement models for stereotyping.
Outcome: The proposed benchmarks lack clarity and assumptions that affect how they conceptualize and operationalize stereotyping.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations