Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing studies have not investigated the differences between different correlation measures in meta-evaluation. |
| Approach: | They analyze 12 common correlation measures using real-world data from six widely-used NLG evaluation datasets and 32 evaluation metrics. |
| Outcome: | The proposed measures exhibit the best performance in discriminative power and ranking consistency . the measures using system-level grouping or Kendall correlation are the least sensitive to score granularity . |
Similar Papers
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics are insufficient to meet requirements for natural language generation. |
| Approach: | They propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities and a method of automatically constructing benchmarks without requiring new human annotations. |
| Outcome: | The proposed framework improves interpretability and provides better performance for 16 representative LLMs. |
Spurious Correlations in Reference-Free Evaluation of Text Generation (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work suggests that reference-free evaluation metrics may rely on spurious correlations with human judgments. |
| Approach: | They propose to use model-based, reference-free evaluation metrics to evaluate natural language generation systems. |
| Outcome: | The proposed metrics achieve high correlations with human judgments, but they may not be robust enough to evaluate their efficacy and robustness. |
Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing definitions of system-level correlations are inconsistent with how they are used to evaluate systems. |
| Approach: | They propose to calculate correlations only on pairs of systems separated by small differences in automatic scores . they propose to use the full test set instead of the subset of summaries judged by humans . |
| Outcome: | The proposed changes improve the accuracy of the estimated correlations on pairs of systems separated by small differences in automatic scores. |
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist (2023.acl-long)
Copied to clipboard
| Challenge: | a systematic review of automatic evaluation metrics for Natural Language Generation (NLG) shows that task-agnostic metrics have a weak correlation with human . |
| Approach: | They propose a framework to assess the effectiveness of automatic metrics in three NLG tasks . they propose task-agnostic and human-aligned metrics to be used for evaluation . |
| Outcome: | The proposed framework provides access to the evaluation tools for three NLG tasks. |
Large-Scale Correlation Analysis of Automated Metrics for Topic Models (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies on topic models lack a correlation between automated coherence metrics and human judgement. |
| Approach: | They propose a sampling approach to mine topics for metric evaluation and extend the analysis to measure topical differences between corpora. |
| Outcome: | The proposed method extends to measure topical differences between corpora and human judgement by using extensive user study. |
Towards a Better Metric for Evaluating Question Generation Systems (D18-1)
Copied to clipboard
| Challenge: | Existing evaluation metrics based on n-gram similarity do not correlate well with human judgments . large datasets for document Question Answering (QA) have enabled the development of end-to-end supervised models . |
| Approach: | They propose a scoring function to capture answerability of questions . they also integrate existing similarity metrics into the scoring function . |
| Outcome: | The proposed scoring function improves human judgments on question answerability . the proposed scoring functions are made publicly available . |
On the Correlation of Word Embedding Evaluation Metrics (2020.lrec-1)
Copied to clipboard
| Challenge: | Word embeddings are geometrical representations of word paradigmatics and syntagmatics. |
| Approach: | They propose to investigate evaluation metrics on various datasets to find correlations . they propose a fast solution to select the best word embeddings among many others . |
| Outcome: | The proposed method could be used to select the best word embeddings among many others. |
Perturbation CheckLists for Evaluating NLG Evaluation Metrics (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output. |
| Approach: | They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation. |
| Outcome: | The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output. |
Is Reference Necessary in the Evaluation of NLG Systems? When and Where? (2024.naacl-long)
Copied to clipboard
| Challenge: | Despite recent advances in reference-free metrics, it has not been well understood when and where they can be used as an alternative to reference-based metrics. |
| Approach: | They propose to use reference-free metrics to evaluate NLG systems . they find they have a higher correlation with human judgment and greater sensitivity to deficiencies in language quality . |
| Outcome: | The proposed metrics exhibit higher correlation with human judgment and greater sensitivity to deficiencies in language quality. |
How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure Evaluation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods to evaluate summary coherence are often evaluated using disparate datasets and metrics. |
| Approach: | They propose to use automatic evaluation to evaluate coherence of summaries by selecting high-scoring candidates. |
| Outcome: | The proposed methods show that they can perform better on an even playing field. |