Rethinking the Agreement in Human Evaluation Tasks (C18-1)

Copied to clipboard

Challenge: In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability.
Approach: They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias.
Outcome: The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks.

Similar Papers

Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement (2020.coling-main)

Copied to clipboard

Challenge: Existing evaluation methods rely on agreement between annotators, which implies a single correct interpretation.
Approach: They propose an agreement-independent quality metric based on answer-coherence to evaluate on expected disagreement.
Outcome: The proposed model shows that agreement is the most important indicator of quality in semantic annotation tasks.
Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)

Copied to clipboard

Challenge: Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer.
Approach: They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines.
Outcome: The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures.
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)

Copied to clipboard

Challenge: a few popular metrics are still used to evaluate language generation systems despite their known limitations.
Approach: They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts .
Outcome: The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set.
What Can We Learn from Collective Human Opinions on Natural Language Inference Data? (2020.emnlp-main)

Copied to clipboard

Challenge: Despite the subjective nature of many NLU evaluations, little attention has been paid to the distribution of human opinions.
Approach: They use a dataset with 464,500 annotations to study Collective HumAn OpinionS . they argue that models lack the ability to recover the distribution over human labels .
Outcome: The proposed dataset examines the distribution of human opinions in NLU evaluation datasets.
The Authenticity Gap in Human Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Using the standard protocol to evaluate NLGs is often violated, resulting in annotator ratings cease to reflect their preferences.
Approach: They propose a human evaluation protocol called system-level probabilistic assessment (SPA) this protocol is based on the assumption that annotators are biased by likert scales .
Outcome: The proposed protocol can recover the ordering of GPT-3 models by size, but less than half of the expected preferences can be recovered when human evaluation is done with the standard protocol.
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
Approach: This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement .
Outcome: This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field.
You Are What You Annotate: Towards Better Models through Annotator Representations (2023.findings-emnlp)

Copied to clipboard

Challenge: Annotator disagreement is ubiquitous in natural language processing tasks.
Approach: They propose to model annotators' idiosyncrasies and account for their idioms by creating representations for each annotator and their annotations.
Outcome: The proposed model improves on an existing dataset with eight annotators with inherent disagreements while increasing model size by 1%.
Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets (D19-1)

Copied to clipboard

Challenge: Having only a few workers generate the majority of dataset examples raises concerns about data diversity .
Approach: They perform a series of experiments to investigate annotator biases in recent NLU datasets . they find that models are able to recognize the most productive annotators .
Outcome: The results show that models can recognize the most productive annotators and do not generalize well to examples from annotator that did not contribute to the training set.
Holistic Inter-Annotator Agreement and Corpus Coherence Estimation in a Large-scale Multilingual Annotation Campaign (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper we examine the complexity of persuasion technique annotation in a multilingual annotation campaign involving 6 languages and approximately 40 annotators.
Approach: They propose a word embedding-based annotator agreement metric and propose 'holistic IAA' metric to measure the coherence of the entire dataset.
Outcome: The proposed method is compared with the existing IAA metrics and its correlation with the results.
Unifying Human and Statistical Evaluation for Natural Language Generation (N19-1)

Copied to clipboard

Challenge: Human evaluation captures quality but fails to capture diversity . statistical evaluation fails to catch models that plagiarize from training set .
Approach: They propose a framework which evaluates both diversity and quality based on the optimal error rate of predicting whether a sentence is human-generated.
Outcome: The proposed framework evaluates diversity and quality on summarization and chit-chat dialogue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations