Challenge: Existing evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability.
Approach: They propose a metric that evaluates natural language generation tasks as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models without training on evaluation datasets.
Outcome: The proposed metric achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which exhibits strong dimension-level / task-level generalization ability and interpretability.

Similar Papers

A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
SESCORE2: Learning Text Generation Evaluation via Synthesizing Realistic Mistakes (2023.acl-long)

Copied to clipboard

Challenge: Existing learned metrics perform unsatisfactory across text generation tasks or require human annotations for training on specific tasks.
Approach: They propose a self-supervised approach to train a model-based metric for text generation evaluation using sentences retrieved from a corpus.
Outcome: The proposed model outperforms all prior unsupervised metrics on four text generation evaluation benchmarks, with an average Kendall improvement of 0.158.
Not All Errors are Equal: Learning Text Generation Metrics using Stratified Error Synthesis (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing learning metrics are limited to tasks where large human ratings are available.
Approach: They propose a model-based natural language generation (NLG) evaluation metric that is highly correlated with human judgements without requiring human annotation.
Outcome: The proposed metric outperforms all prior unsupervised metrics on multiple NLG tasks including translation, image captioning, and WebNLG text generation.
Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Natural language generation (NLG) tasks have complex nature and require manual evaluation.
Approach: They propose a unifying perspective based on the nature of information change in NLG tasks . they propose 'information alignment' metrics that can be used to evaluate different aspects of NLG .
Outcome: The proposed metrics achieve stronger or comparable correlations with human judgement compared to state-of-the-art metrics in diverse tasks.
Toward Human-Like Evaluation for Natural Language Generation with Error Analysis (2023.acl-long)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) have been used to evaluate language generation tasks . pretrained error analysis can be used to refine the generated sentence toward higher confidence .
Approach: They propose to combine pretrained language model based metrics with human-like error analysis to improve sentence confidence.
Outcome: The proposed method outperforms top-scoring metrics in 19/25 settings.
RoMe: A Robust Metric for Evaluating Natural Language Generation (2022.acl-long)

Copied to clipboard

Challenge: Empirical results suggest that RoMe has a stronger correlation to human judgment over state-of-the-art metrics in evaluating system-generated sentences across several NLG tasks.
Approach: They propose an automatic evaluation metric incorporating several core aspects of natural language understanding (language competence, syntactic and semantic variation).
Outcome: The proposed evaluation metric is trained on language features such as semantic similarity combined with tree edit distance and grammatical acceptability, using a self-supervised neural network.
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications (2022.naacl-main)

Copied to clipboard

Challenge: Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text.
Approach: They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation.
Outcome: The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations.
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) produce incomplete or selectively omit key information . omissions of key information or misrepresentation of conflicting evidence can cause harm .
Approach: They propose a method that decomposes texts into atomic statements and uses natural language inference to identify missing facts and a Q A-based metric that extracts question-answer pairs and compares responses across sources.
Outcome: The proposed evaluation metrics show they perform better than more complex metrics, but at a cost.
Towards a Better Metric for Evaluating Question Generation Systems (D18-1)

Copied to clipboard

Challenge: Existing evaluation metrics based on n-gram similarity do not correlate well with human judgments . large datasets for document Question Answering (QA) have enabled the development of end-to-end supervised models .
Approach: They propose a scoring function to capture answerability of questions . they also integrate existing similarity metrics into the scoring function .
Outcome: The proposed scoring function improves human judgments on question answerability . the proposed scoring functions are made publicly available .
NLG-Metricverse: An End-to-End Library for Evaluating Natural Language Generation (2022.coling-1)

Copied to clipboard

Challenge: Natural language generation models are a key component of deep learning, says aaron eliott . he says it is crucial to develop and apply better metrics for NLG evaluation .
Approach: a new open-source library for NLG evaluation is created to facilitate researchers to judge the effectiveness of their models. the framework provides a living collection of NLG metrics in a unified and easy-to-use environment.
Outcome: a new open-source library for NLG evaluation aims to improve performance of models . the framework provides tools to apply, analyze, compare, and visualize the metrics .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations