Challenge: ML benchmarks have been criticized for their construct validity, fragility of the design and task choices.
Approach: They propose a framework for ranking systems in multi-task benchmarks under the principles of the social choice theory and propose 'vote'n'rank' procedures are more robust than the mean average while being able to handle missing performance scores and determine conditions under which the system becomes the winner.
Outcome: The proposed framework can be utilised to draw new insights on benchmarking in several ML sub-fields and identify the best-performing systems in research and development case studies.

Similar Papers

A Study of Latent Structured Prediction Approaches to Passage Reranking (N19-1)

Copied to clipboard

Challenge: a structured output framework is useful for learning to rank problems . current approaches for answer sentence reranking are mostly based on pairwise ranking signals or simple binary classification.
Approach: They propose a structured output approach which regards rankings as latent variables . they propose an inference procedure to find the max-violating ranking based on decomposition of the corresponding loss.
Outcome: The proposed approach solves the optimization problem on WikiQA and TREC13 datasets.
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference (2025.findings-naacl)

Copied to clipboard

Challenge: Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences.
Approach: They propose a system-level evaluation framework that ranks LLMs based on their alignment with human preferences.
Outcome: The proposed framework aims to rank LLMs based on their performance and alignment with human preferences.
JuStRank: Benchmarking LLM Judges for System Ranking (2025.acl-long)

Copied to clipboard

Challenge: Recent work has focused on instance-based evaluation of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.
Approach: They propose to validate the quality of the LLM judge itself by comparing system scores to a human-based ranking.
Outcome: The proposed model fails to validate the quality of the judge itself, ignoring critical factors affecting system-level ranking, such as a judge’s positive or negative bias towards certain systems.
Reproducible and Efficient Benchmarks for Hyperparameter Optimization of Neural Machine Translation Systems (2020.tacl-1)

Copied to clipboard

Challenge: Optimal versus suboptimal hyperparameters can lead to dramatic swings in system performance.
Approach: They propose to use a library of pre-trained models for fast, low cost HPO experimentation and to propose metrics for evaluating HPO methods on NMT.
Outcome: The proposed method uses a library of pre-trained models for fast, low cost experimentation on neural machine translation (NMT) .
The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks (2023.acl-short)

Copied to clipboard

Challenge: omnipresence of large pre-trained language models has fueled concerns regarding systematic biases carried over from underlying data into the applications they are used in.
Approach: They propose to compare social biases with non-social biase masked by alternate constructions that maintain the essence of their social bias.
Outcome: The proposed benchmarks underestimate or overestimate the social bias in a given model.
Boosting Self-Consistency with Ranking (2026.acl-srw)

Copied to clipboard

Challenge: Existing approaches to improve performance of large language models include self-consistency, RISC, extended reasoning, and iterative self-correction.
Approach: They propose a test-time scaling technique that uses multiple features to score candidate answers in self-consistency as a ranking problem.
Outcome: The proposed method achieves better accuracy-efficiency trade-off than standard self-consistency and strong baselines on three datasets.
Towards a Better Metric for Evaluating Question Generation Systems (D18-1)

Copied to clipboard

Challenge: Existing evaluation metrics based on n-gram similarity do not correlate well with human judgments . large datasets for document Question Answering (QA) have enabled the development of end-to-end supervised models .
Approach: They propose a scoring function to capture answerability of questions . they also integrate existing similarity metrics into the scoring function .
Outcome: The proposed scoring function improves human judgments on question answerability . the proposed scoring functions are made publicly available .
What Will it Take to Fix Benchmarking in Natural Language Understanding? (2021.naacl-main)

Copied to clipboard

Challenge: Evaluation for many natural language understanding (NLU) tasks is broken due to unreliable and biased systems scoring so high on standard benchmarks.
Approach: They argue that current benchmarks fail at four criteria for evaluation . they argue that adversarial data collection does not address the causes of failures .
Outcome: The proposed frameworks fail at four criteria, and adversarial data collection does not address the causes of these failures, the authors argue . restoring a healthy evaluation ecosystem will require significant progress in the design of benchmark datasets, reliability with which they are annotated, their size, and the ways they handle social bias.
Beyond One-Size-Fits-All: Tailored Benchmarks for Efficient Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Existing efficient methods estimate performance of models on large benchmarks, but these methods rely on the assumption that target models have high prediction consistency with source models.
Approach: They propose a method that conducts customized evaluation tailored to each target model.
Outcome: The proposed method reduces the MAE of estimates by 31.4% on benchmarks across 300 models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations