Mark-Evaluate: Assessing Language Generation using Population Estimation Methods (2020.coling-main)
Copied to clipboard
| Challenge: | Existing population estimation methods focus on open populations or closed populations, but our methods show a higher correlation to human evaluation than existing metrics on several challenging tasks. |
| Approach: | They propose a family of metrics to assess language generation derived from population estimation methods widely used in ecology. |
| Outcome: | The proposed methods show a higher correlation to human evaluation than existing metrics on several challenging tasks, namely unconditional language generation, machine translation, and text summarization. |
Similar Papers
Unifying Human and Statistical Evaluation for Natural Language Generation (N19-1)
Copied to clipboard
| Challenge: | Human evaluation captures quality but fails to capture diversity . statistical evaluation fails to catch models that plagiarize from training set . |
| Approach: | They propose a framework which evaluates both diversity and quality based on the optimal error rate of predicting whether a sentence is human-generated. |
| Outcome: | The proposed framework evaluates diversity and quality on summarization and chit-chat dialogue. |
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)
Copied to clipboard
| Challenge: | a few popular metrics are still used to evaluate language generation systems despite their known limitations. |
| Approach: | They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts . |
| Outcome: | The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set. |
Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation (D19-1)
Copied to clipboard
| Challenge: | Existing evaluation methods for natural language generation are inadequate . distinguishing machine-generated text is challenging even for human evaluators . |
| Approach: | They compare human-based evaluators with automated evaluation procedures . they find human evaluers do not correlate well with discriminative evalators . |
| Outcome: | The proposed evaluation methods are compared with a dozen state-of-the-art generators for online product reviews. |
Evaluating the Evaluation of Diversity in Natural Language Generation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for controlling diversity by tuning a “decoding parameter” affect form but not meaning. |
| Approach: | They propose a framework that measures correlation between a diversity metric and a parameter that controls some aspect of diversity in generated text. |
| Outcome: | The proposed framework outperforms existing methods in estimating diversity . it shows that humans outperformed existing methods but affect form but not meaning . |
Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand (2022.naacl-main)
Copied to clipboard
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, Noah A. Smith
| Challenge: | Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models . |
| Approach: | They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation. |
| Outcome: | The proposed leaderboards track progress in language generation models and metrics for their evaluation. |
On the Effectiveness of Automated Metrics for Text Generation Systems (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation methods lack a sound theoretical foundation for evaluation campaigns . imperfect automated metrics and insufficiently sized test sets are some of the factors that cause uncertainty. |
| Approach: | They propose a theoretical framework that incorporates different sources of uncertainty, such as imperfect automated metrics and insufficiently sized test sets. |
| Outcome: | The proposed model can be leveraged to improve evaluation protocols regarding reliability, robustness, and significance of the evaluation outcome. |
Evaluating the Evaluation of Diversity in Commonsense Generation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics for commonsense generation are unclear on which metrics are best suited for evaluating the diversity of outputs. |
| Approach: | They propose to use a large language model to analyze commonsense generation data to determine which diversity metrics are best suited for commonsensing. |
| Outcome: | The proposed metrics outperform form-based metrics and show high correlations with the LLM-based ratings. |
ICE-Score: Instructing Large Language Models to Evaluate Code (2024.findings-eacl)
Copied to clipboard
| Challenge: | Recent advances in the field of natural language generation have facilitated the use of large language models to assess the quality of generated text. |
| Approach: | They propose a new evaluation metric by instructing large language models for code assessments using a set of programming languages. |
| Outcome: | The proposed evaluation metric surpasses state-of-the-art metrics for code generation, delivering high levels of accuracy and consistency across programming languages and tasks. |
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)
Copied to clipboard
| Challenge: | n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear. |
| Approach: | They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics. |
| Outcome: | The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand. |