| Challenge: | Existing methods to evaluate natural language systems are expensive and expensive. |
| Approach: | They propose to combine automatic metrics with human judgment to obtain an unbiased estimator at lower cost than human evaluation alone. |
| Outcome: | The proposed estimator reduces the cost of evaluating summarization and open-response questions by 7-13%. |
Similar Papers
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for judging metrics are sensitive to the translations used for evaluation, leading to falsely confident conclusions about a metric’s efficacy. |
| Approach: | They propose a method for thresholding performance improvement under an automatic metric against human judgements by using a pairwise system ranking method. |
| Outcome: | The proposed method allows quantification of type I versus type II errors incurred, i.e., insignificant human differences in system quality that are accepted, and significant human differences that are rejected. |
The statistical advantage of automatic NLG metrics at the system level (2021.acl-long)
Copied to clipboard
| Challenge: | Statistically, humans are unbiased, high variance estimators, while metrics are biased, low variance estimator. |
| Approach: | They compare automatic metrics to humans and a derived, perfect segment-level annotator by applying a bias-variance-noise decomposition to adjust the error to a noise-free, infinite test set setting. |
| Outcome: | The proposed method outperforms humans and a derived, perfect segment-level annotator in two settings. |
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)
Copied to clipboard
| Challenge: | a few popular metrics are still used to evaluate language generation systems despite their known limitations. |
| Approach: | They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts . |
| Outcome: | The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set. |
Automating Human Evaluation of Dialogue Systems (2022.naacl-srw)
Copied to clipboard
| Challenge: | a recent study shows that human evaluations of dialogue systems weakly reflect human judgments. |
| Approach: | They propose a BERT-based model that fine-tunes a model with three prediction heads to predict whether the system-generated output is natural, fluent, and informative. |
| Outcome: | The proposed model achieves an average accuracy of 77% over the 3 labels . it also uses three different models to compute the labels compared to three separate models . |
Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation (2023.acl-long)
Copied to clipboard
Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev
| Challenge: | Existing studies for summarization evaluation exhibit low inter-annotator agreement or lack scale. |
| Approach: | They propose a modified summarization salience protocol based on fine-grained semantic units and a robust summarizing evaluation benchmark. |
| Outcome: | The proposed protocol is based on fine-grained semantic units and allows for high inter-annotator agreement. |
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? (2025.acl-short)
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics are based on procedures that diverge from human evaluation. |
| Approach: | They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system. |
| Outcome: | The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics. |
Rethinking the Agreement in Human Evaluation Tasks (C18-1)
Copied to clipboard
| Challenge: | In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability. |
| Approach: | They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias. |
| Outcome: | The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks. |
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
A Measure of the System Dependence of Automated Metrics (2025.acl-short)
Copied to clipboard
| Challenge: | Recent advances in machine translation evaluations are expensive and time-intensive. |
| Approach: | They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently. |
| Outcome: | The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure. |
On the Limitations of Reference-Free Evaluations of Generated Text (2022.emnlp-main)
Copied to clipboard
| Challenge: | a recent study has shown that evaluation metrics which accurately estimate the quality of generated text are limited in their ability to evaluate generated text. |
| Approach: | They argue that reference-free metrics are limited in their ability to evaluate generated text . they recommend that they be used as diagnostic tools for analyzing and understanding model behavior . |
| Outcome: | The proposed evaluation metrics are limited in their ability to evaluate generated text . they can be optimized at test time, can be biased against models with similar outputs . |