Challenge: Existing evaluation methods rely on agreement between annotators, which implies a single correct interpretation.
Approach: They propose an agreement-independent quality metric based on answer-coherence to evaluate on expected disagreement.
Outcome: The proposed model shows that agreement is the most important indicator of quality in semantic annotation tasks.

Similar Papers

Rethinking the Agreement in Human Evaluation Tasks (C18-1)

Copied to clipboard

Challenge: In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability.
Approach: They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias.
Outcome: The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks.
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)

Copied to clipboard

Challenge: supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data.
Approach: They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity.
Outcome: The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness.
Validity, Agreement, Consensuality and Annotated Data Quality (2022.lrec-1)

Copied to clipboard

Challenge: a wide consensus is rife regarding the need for reference annotated datasets . however, the creation of such datasets is accompanied by theorectical and practical issues .
Approach: They propose to use agreement among annotators as an indicator of consensus . they argue that it is difficult to produce gold-standard annotated datasets .
Outcome: The proposed model focuses on the complex relations between agreement and reference and the emergence of consensus.
A Crowdsourced Frame Disambiguation Corpus with Ambiguity (N19-1)

Copied to clipboard

Challenge: Using crowdsourcing, we have found that inter-annotator disagreement is at least partly caused by ambiguity inherent to the text and frames.
Approach: They propose a crowdsourcing approach to capture inter-annotator disagreement by a list of frames with disagreement-based scores that express the confidence with which each frame applies to the word.
Outcome: The proposed approach captures disagreement between the annotations of 1,000 word-sentence pairs and scores on the likelihood that each frame applies to the word.
You Are What You Annotate: Towards Better Models through Annotator Representations (2023.findings-emnlp)

Copied to clipboard

Challenge: Annotator disagreement is ubiquitous in natural language processing tasks.
Approach: They propose to model annotators' idiosyncrasies and account for their idioms by creating representations for each annotator and their annotations.
Outcome: The proposed model improves on an existing dataset with eight annotators with inherent disagreements while increasing model size by 1%.
Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning (2023.acl-long)

Copied to clipboard

Challenge: Annotator disagreements are resolved before learning takes place, but researchers question the performance of a system when annotators disagree.
Approach: They propose a method that uses language features and label distributions to pool similar items into larger labels.
Outcome: The proposed method is based on five publicly available datasets with varying levels of disagreements on social media and in the wild using a dataset from Facebook.
Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)

Copied to clipboard

Challenge: Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer.
Approach: They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines.
Outcome: The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures.
What Can We Learn from Collective Human Opinions on Natural Language Inference Data? (2020.emnlp-main)

Copied to clipboard

Challenge: Despite the subjective nature of many NLU evaluations, little attention has been paid to the distribution of human opinions.
Approach: They use a dataset with 464,500 annotations to study Collective HumAn OpinionS . they argue that models lack the ability to recover the distribution over human labels .
Outcome: The proposed dataset examines the distribution of human opinions in NLU evaluation datasets.
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values.
Approach: They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data.
Outcome: The proposed method breaks down disagreements by asking raters how they think others would annotate the data.
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to assess annotator judgments aggregate disagreement into a single gold label . et al., 2022) show disagreement is diffuse, but standard approaches are not rigorous .
Approach: They propose a schema-level diagnostic for auditing expert-designed annotation schemas prior to gold-label commitment . they find disagreement is not diffuse: instability concentrates in a few criteria, while nearly half of covered sentences activate multiple categories.
Outcome: The proposed diagnostic separates unstable criteria with hard-to-operationalize boundaries and systematic overlap that blurs the boundaries between mutually exclusive categories.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations