HAUSER: Towards Holistic and Automatic Evaluation of Simile Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Similes are a crucial part of creative writing, but there is still a lack of evaluation metrics for simile generation. |
| Approach: | They propose to use similes as a tool to evaluate simile generation metrics . they propose to combine five criteria and automatic metrics for each criterion . |
| Outcome: | The proposed metrics are significantly more correlated with human ratings from each perspective compared with prior automatic metrics. |
Similar Papers
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies on automatic story generation (ASG) rely on human criteria, but there is little research on how well they correlate with human criteria. |
| Approach: | They propose to use human criteria to evaluate automatic story generation (ASG) their paper proposes to use HANNA to quantitatively evaluate correlations between 72 automatic metrics and human criteria. |
| Outcome: | The proposed model compared human criteria with automatic criteria and found that they were significantly better than human criteria. |
GRUEN for Evaluating Linguistic Quality of Generated Text (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation metrics focus on content selection, not linguistic quality . proposed GRUEN measures Grammaticality, non-redundancy, focUs, structure and coherence of generated text. |
| Approach: | They propose to use a BERT-based model and a class of syntactic, semantic, and contextual features to examine the system output. |
| Outcome: | Experiments show that the proposed metric correlates highly with human judgments. |
Towards Holistic and Automatic Evaluation of Open-Domain Dialogue Generation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient. |
| Approach: | They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics. |
| Outcome: | The proposed metrics show strong correlations with human judgments. |
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)
Copied to clipboard
| Challenge: | a few popular metrics are still used to evaluate language generation systems despite their known limitations. |
| Approach: | They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts . |
| Outcome: | The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set. |
Perturbation CheckLists for Evaluating NLG Evaluation Metrics (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output. |
| Approach: | They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation. |
| Outcome: | The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output. |
NLG-Metricverse: An End-to-End Library for Evaluating Natural Language Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Natural language generation models are a key component of deep learning, says aaron eliott . he says it is crucial to develop and apply better metrics for NLG evaluation . |
| Approach: | a new open-source library for NLG evaluation is created to facilitate researchers to judge the effectiveness of their models. the framework provides a living collection of NLG metrics in a unified and easy-to-use environment. |
| Outcome: | a new open-source library for NLG evaluation aims to improve performance of models . the framework provides tools to apply, analyze, compare, and visualize the metrics . |
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement. |
| Approach: | They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction. |
| Outcome: | The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics. |
Unifying Human and Statistical Evaluation for Natural Language Generation (N19-1)
Copied to clipboard
| Challenge: | Human evaluation captures quality but fails to capture diversity . statistical evaluation fails to catch models that plagiarize from training set . |
| Approach: | They propose a framework which evaluates both diversity and quality based on the optimal error rate of predicting whether a sentence is human-generated. |
| Outcome: | The proposed framework evaluates diversity and quality on summarization and chit-chat dialogue. |
A Human Evaluation of AMR-to-English Generation Systems (2020.coling-main)
Copied to clipboard
| Challenge: | a recent human evaluation of AMR generation systems is compared to automated metrics. |
| Approach: | They propose a human evaluation which collects fluency and adequacy scores and categorization of error types for AMR generation systems. |
| Outcome: | The results show that human evaluations are more nuanced than automated metrics. |