Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations do not evaluate the same aspect of quality, resulting in unclear comparability and low repeatability. |
| Approach: | They propose to use a standard set of qualitycriterion names and definitions to establish comparability of existing evaluations. |
| Outcome: | The proposed taxonomy combines 114 quality criteria from 3 surveys of 933 evaluations in NLP and is used to establish comparability of existing evaluations and guide the design of new evaluations. |
Similar Papers
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability. |
| Approach: | They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria. |
| Outcome: | The proposed system is based on 11 common aspects with different evaluation criteria. |
Re-Examining Summarization Evaluation across Multiple Quality Criteria (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a number of automated evaluation metrics are evaluated by multiple quality criteria, such as relevance, consistency, fluency and coherence. |
| Approach: | They propose a method that removes the confounding variable and detects unreliable correlations. |
| Outcome: | The proposed method detects unreliable correlations between QCs and human scores . it is based on a multi-QC setup, but it fails to detect summary corruptions . |
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications (2022.naacl-main)
Copied to clipboard
| Challenge: | Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text. |
| Approach: | They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation. |
| Outcome: | The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations. |
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
Is NLP Ready for Standardization? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | a number of scientific fields, including telecommunications, networks and multimedia, lack standards in the field of NLP. |
| Approach: | They propose to examine how NLP lacks standards and how that can impact society, industry and regulations. |
| Outcome: | The proposed standards examine the needs of NLP researchers and industry . they argue that the lack of standards can impact the field, society and industry. |
Fair Enough: Standardizing Evaluation and Model Selection for Fairness Research in NLP (2023.eacl-main)
Copied to clipboard
| Challenge: | Modern NLP systems exhibit a range of biases, which a growing literature on model debiasing attempts to correct. |
| Approach: | They propose to clarify the current situation and plot a course for meaningful progress in fair learning by making clear inter-relations among the current gamut of methods and their relation to fairness theory. |
| Outcome: | The proposed approach addresses the practical problem of model selection, which involves a trade-off between fairness and accuracy and has led to systemic issues in fairness research. |
Perturbation CheckLists for Evaluating NLG Evaluation Metrics (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation are inadequate . existing metrics are not robust against simple perturbations and disagree with scores assigned by humans to perturbed output. |
| Approach: | They propose to propose checks which perturb the output and target a specific criteria and then use them to refine their evaluation. |
| Outcome: | The proposed templates show that existing evaluation metrics are not robust against simple perturbations and disagree with human scores on the perturbed output. |
Reference-Free Evaluation of Taxonomies (2026.findings-acl)
Copied to clipboard
| Challenge: | Taxonomies are used to classify items, ideas or organisms based on shared characteristics. |
| Approach: | They introduce two reference-free metrics for quality evaluation of taxonomies in the absence of labels. |
| Outcome: | The proposed metrics correlate well with F1 against ground truth taxonomies on five taxonomies and improve hierarchical classification when used with label hierarchies. |
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics are insufficient to meet requirements for natural language generation. |
| Approach: | They propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities and a method of automatically constructing benchmarks without requiring new human annotations. |
| Outcome: | The proposed framework improves interpretability and provides better performance for 16 representative LLMs. |
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)
Copied to clipboard
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
| Challenge: | Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution. |
| Approach: | They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage. |
| Outcome: | The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage. |