Challenge: Existing studies do not consider variance change due to metric model errors, which can lead to wrong conclusions.
Approach: They establish the mathematical foundation of significance testing for model-based metrics . they show that metric errors can change the conclusions in certain experiments .
Outcome: The proposed method can be used to derive accurate conclusions using model evaluations.

Similar Papers

The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing (P18-1)

Copied to clipboard

Challenge: Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental.
Approach: They propose a protocol for statistical significance test selection in NLP setups . they propose he proposes a survey of the most relevant tests to help guide the protocol .
Outcome: The proposed protocol includes a survey of the most relevant tests.
Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics are conflated and can mislead models, resulting in downstream harms.
Approach: They propose a framework for conceptualizing and evaluating the reliability and validity of evaluation metrics based on empirical data.
Outcome: The proposed framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data.
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for estimating model reliability are based on a few output responses per item.
Approach: They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing.
Outcome: The proposed method can help researchers make better decisions about how to collect data for AI evaluation.
Measuring Robustness for NLP (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to evaluate NLP models are limited to news domains and cannot be generalized to other domains.
Approach: They propose a measure of NLP quality based on robustness . they measure consistency of cross-domain accuracy and introduce coefficient of variation and gamma-Robustness based upon human evaluation .
Outcome: The proposed approach shows higher agreement with human evaluation than accuracy scores on ranking machine translation systems.
Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess Hypotheses (2020.acl-main)

Copied to clipboard

Challenge: Empirical research in natural language processing has adopted a narrow set of principles for assessing hypotheses . alternative approaches to assess hypothese rely on p-value computation, which suffers from several known issues.
Approach: They propose to compare different methods for assessing hypotheses . they argue that practitioners should first decide their target hypothesis before choosing a method .
Outcome: The proposed method differs from other methods, but is not widely used in NLP . the proposed method is based on a p-value computation, but has a small gap in accuracy .
Towards Robust Comparisons of NLP Models: A Case Study (2025.coling-main)

Copied to clipboard

Challenge: Existing statistical tests to compare the test scores of different NLP models have been proposed to account for nuisance factors such as noise, randomness, or hyperparameter values.
Approach: They propose a regression analysis which isolates the effect of nuisance factors from the effects of the models’ capabilities.
Outcome: The proposed model is able to show that the difference between BioLinkBERT and MSR BiomedBERT is 7 times smaller than previously reported.
Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa (2022.acl-demo)

Copied to clipboard

Challenge: Developing better methods for a task is a common feature of the computational linguistics literature.
Approach: They propose to use bootstrap to compute significance levels with the BOOtSTrap SAmpling procedure to evaluate models that predict hard labels and soft labels as well.
Outcome: The proposed method can be used to evaluate models that predict hard labels and soft labels on benchmark data sets.
Evaluating and Characterizing Human Rationales (2020.emnlp-main)

Copied to clipboard

Challenge: a new study examines how human rationales perform on automatic metrics . human-generated rationale evaluation is difficult because of its ambiguity .
Approach: They propose to use model-dependent baseline performance to evaluate rationale quality . they propose to also use "fidelity curves" to reveal properties such as irrelevance and redundancy .
Outcome: The proposed methods characterize rationale quality based on model retraining and using "fidelity curves" the proposed methods lead to actionable suggestions for evaluating and characterizing rationales .
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)

Copied to clipboard

Challenge: a few popular metrics are still used to evaluate language generation systems despite their known limitations.
Approach: They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts .
Outcome: The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set.
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios (2024.findings-emnlp)

Copied to clipboard

Challenge: Using large language models, we evaluated their robustness on multiple datasets.
Approach: They propose a new metric for assessing model robustness by empirical evaluation of several models on multiple datasets.
Outcome: The proposed metric is based on a set of datasets that are constructed by introducing naturally-occurring, non-malicious perturbations or by generating semantically equivalent paraphrases of input questions or statements.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations