Challenge: Using standard metrics in the presence of poor labels masks label and model quality . evaluation techniques accounting for unreliable labels reveal important flaws, including spurious correlations and nonrandom racial biases .
Approach: They analyze human labels, GPT model ratings, and transformer encoder model ratings . they show that standard metrics in the presence of poor labels mask label and model quality .
Outcome: The proposed methods mask label and model quality even in the presence of poor models.

Similar Papers

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.
Approach: They propose to use an ensemble of large language models to flag mislabeled examples by using an LLM-as-a-judge approach to detect label errors in existing datasets.
Outcome: The proposed method improves label accuracy and consistency in large language models.
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction (2024.acl-long)

Copied to clipboard

Challenge: Current benchmarks for social biases have limitations in scope, grounding, quality and human effort required.
Approach: They propose to use a language model to help with the development of bias benchmarks . they extend previous work to a new community and set of biases: the Jewish community and antisemitism .
Outcome: The proposed LLM does not perform well on the Jewish community and antisemitism task.
Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human Evaluation (2023.acl-long)

Copied to clipboard

Challenge: Existing studies for summarization evaluation exhibit low inter-annotator agreement or lack scale.
Approach: They propose a modified summarization salience protocol based on fine-grained semantic units and a robust summarizing evaluation benchmark.
Outcome: The proposed protocol is based on fine-grained semantic units and allows for high inter-annotator agreement.
Evaluating Large Language Models on Wikipedia-Style Survey Generation (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that large language models can perform well in general tasks, but their effectiveness and limitations in domainspecific tasks remain unclear.
Approach: They examine the proficiency of Large Language Models (LLMs) in generating succinct survey articles specific to the niche field of NLP in computer science.
Outcome: The LLMs perform better in generating succinct survey articles specific to the niche field of NLP in computer science, compared to human-authored surveys, but they exhibit bias in evaluation.
All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text (2021.acl-long)

Copied to clipboard

Challenge: evaluators distinguish between human- and machine-authored text in three domains without training . evals' accuracy improved up to 55%, but it did not significantly improve across the three domain.
Approach: They examine the role untrained human evaluations play in NLG evaluation and propose ways to improve their evaluations.
Outcome: The evaluators distinguished between human- and machine-authored text at random chance level without training, but their accuracy did not improve across the three domains.
The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: a paper argues that human label variation impacts all stages of the ML pipeline . human label variations are often considered noise due to disagreement, subjectivity in annotation or multiple plausible answers.
Approach: They propose to reconcile different notions of human label variation and propose a repository of publicly-available datasets with un-aggregated labels.
Outcome: The proposed approaches are compared with publicly available datasets with un-aggregated labels and identify gaps.
VariErr NLI: Separating Annotation Error from Human Label Variation (2024.acl-long)

Copied to clipboard

Challenge: Existing work on label variation and annotation errors has focused on them in isolation.
Approach: They propose a 2-round annotation procedure to separate human label variation from annotation errors by pairing valid explanations with annotators' validations.
Outcome: The proposed procedure is based on the NLI task in English and contains 7,732 valid judgements on 1,933 explanations for 500 re-annotated items.
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for estimating model reliability are based on a few output responses per item.
Approach: They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing.
Outcome: The proposed method can help researchers make better decisions about how to collect data for AI evaluation.
Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to model complex subjective tasks in natural language are limited by significant variation in annotations.
Approach: They propose a simple in-context learning binary filtering baseline that estimates the reasonableness of a document-label pair.
Outcome: The proposed approach can be integrated into annotation pipelines to enhance signal-to-noise ratios.
Refined Assessment for Translation Evaluation: Rethinking Machine Translation Evaluation in the Era of Human-Level Systems (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, traditional evaluation methods struggle to detect subtle translation errors.
Approach: They propose to use a dataset of human evaluations for English–Russian translations created by professional linguists to enable consistent and rich annotation.
Outcome: The proposed protocol allows expert assessments without time pressure to yield substantially different results from standard evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations