Challenge: Generative AI systems are becoming ubiquitous for all kinds of modalities . evaluation of generated outputs is increasingly difficult due to cost and complexity of human evaluations.
Approach: They propose to evaluate preference ratings on sign accuracy and favoritism . they propose to use automated metrics to assess generated outputs .
Outcome: The proposed evaluations of preference ratings rely on correlation to human judgments or sign accuracy scores, but this does not tell the whole story.

Similar Papers

Correction of Errors in Preference Ratings from Automated Metrics for Text Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods are over-confident in assigning significant differences between systems . Currently, the most reliable evaluation methods for text generation are human-based evaluations.
Approach: They propose to combine human ratings with automated ratings to reduce the amount of human ratings needed to arrive at robust results.
Outcome: The proposed evaluation protocol reduces the amount of human ratings by 50% while yielding the same evaluation outcome as the pure human evaluation in 95% of cases.
JuStRank: Benchmarking LLM Judges for System Ranking (2025.acl-long)

Copied to clipboard

Challenge: Recent work has focused on instance-based evaluation of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.
Approach: They propose to validate the quality of the LLM judge itself by comparing system scores to a human-based ranking.
Outcome: The proposed model fails to validate the quality of the judge itself, ignoring critical factors affecting system-level ranking, such as a judge’s positive or negative bias towards certain systems.
Towards a Better Metric for Evaluating Question Generation Systems (D18-1)

Copied to clipboard

Challenge: Existing evaluation metrics based on n-gram similarity do not correlate well with human judgments . large datasets for document Question Answering (QA) have enabled the development of end-to-end supervised models .
Approach: They propose a scoring function to capture answerability of questions . they also integrate existing similarity metrics into the scoring function .
Outcome: The proposed scoring function improves human judgments on question answerability . the proposed scoring functions are made publicly available .
R-Fairness: Assessing Fairness of Ranking in Subjective Data (2025.acl-long)

Copied to clipboard

Challenge: Subjective data, reflecting individual opinions, permeates platforms like Yelp and Amazon . despite the prevalence of such platforms, little attention has been given to fairness in their context .
Approach: They propose a fairness assessment pipeline that starts with data collection phase and then iterates through rated items.
Outcome: The proposed approach favors groups writing best-ranked reviews over others on collaborative rating platforms.
A Measure of the System Dependence of Automated Metrics (2025.acl-short)

Copied to clipboard

Challenge: Recent advances in machine translation evaluations are expensive and time-intensive.
Approach: They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently.
Outcome: The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure.
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)

Copied to clipboard

Challenge: Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate.
Approach: They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
Outcome: The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation (2026.acl-srw)

Copied to clipboard

Challenge: Existing evaluation metrics for natural language generation are expensive and time-consuming.
Approach: They propose a framework that utilizes LLMs to generate synthetic evaluation datasets . they propose meta-correlation to measure alignment between metric rankings and human benchmarks based on synthetic data .
Outcome: The proposed framework achieves meta-correlations exceeding 0.9 in multilingual QA and replaces human judgment with synthetic evaluation datasets.
Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation (D19-1)

Copied to clipboard

Challenge: Existing evaluation methods for natural language generation are inadequate . distinguishing machine-generated text is challenging even for human evaluators .
Approach: They compare human-based evaluators with automated evaluation procedures . they find human evaluers do not correlate well with discriminative evalators .
Outcome: The proposed evaluation methods are compared with a dozen state-of-the-art generators for online product reviews.
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist (2023.acl-long)

Copied to clipboard

Challenge: a systematic review of automatic evaluation metrics for Natural Language Generation (NLG) shows that task-agnostic metrics have a weak correlation with human .
Approach: They propose a framework to assess the effectiveness of automatic metrics in three NLG tasks . they propose task-agnostic and human-aligned metrics to be used for evaluation .
Outcome: The proposed framework provides access to the evaluation tools for three NLG tasks.
The statistical advantage of automatic NLG metrics at the system level (2021.acl-long)

Copied to clipboard

Challenge: Statistically, humans are unbiased, high variance estimators, while metrics are biased, low variance estimator.
Approach: They compare automatic metrics to humans and a derived, perfect segment-level annotator by applying a bias-variance-noise decomposition to adjust the error to a noise-free, infinite test set setting.
Outcome: The proposed method outperforms humans and a derived, perfect segment-level annotator in two settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations