Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement (2020.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation methods rely on agreement between annotators, which implies a single correct interpretation. |
| Approach: | They propose an agreement-independent quality metric based on answer-coherence to evaluate on expected disagreement. |
| Outcome: | The proposed model shows that agreement is the most important indicator of quality in semantic annotation tasks. |
Similar Papers
Rethinking the Agreement in Human Evaluation Tasks (C18-1)
Copied to clipboard
| Challenge: | In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability. |
| Approach: | They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias. |
| Outcome: | The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks. |
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
Validity, Agreement, Consensuality and Annotated Data Quality (2022.lrec-1)
Copied to clipboard
| Challenge: | a wide consensus is rife regarding the need for reference annotated datasets . however, the creation of such datasets is accompanied by theorectical and practical issues . |
| Approach: | They propose to use agreement among annotators as an indicator of consensus . they argue that it is difficult to produce gold-standard annotated datasets . |
| Outcome: | The proposed model focuses on the complex relations between agreement and reference and the emergence of consensus. |
A Crowdsourced Frame Disambiguation Corpus with Ambiguity (N19-1)
Copied to clipboard
| Challenge: | Using crowdsourcing, we have found that inter-annotator disagreement is at least partly caused by ambiguity inherent to the text and frames. |
| Approach: | They propose a crowdsourcing approach to capture inter-annotator disagreement by a list of frames with disagreement-based scores that express the confidence with which each frame applies to the word. |
| Outcome: | The proposed approach captures disagreement between the annotations of 1,000 word-sentence pairs and scores on the likelihood that each frame applies to the word. |
You Are What You Annotate: Towards Better Models through Annotator Representations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Annotator disagreement is ubiquitous in natural language processing tasks. |
| Approach: | They propose to model annotators' idiosyncrasies and account for their idioms by creating representations for each annotator and their annotations. |
| Outcome: | The proposed model improves on an existing dataset with eight annotators with inherent disagreements while increasing model size by 1%. |
Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Annotator disagreements are resolved before learning takes place, but researchers question the performance of a system when annotators disagree. |
| Approach: | They propose a method that uses language features and label distributions to pool similar items into larger labels. |
| Outcome: | The proposed method is based on five publicly available datasets with varying levels of disagreements on social media and in the wild using a dataset from Facebook. |
Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)
Copied to clipboard
| Challenge: | Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer. |
| Approach: | They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines. |
| Outcome: | The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures. |
What Can We Learn from Collective Human Opinions on Natural Language Inference Data? (2020.emnlp-main)
Copied to clipboard
| Challenge: | Despite the subjective nature of many NLU evaluations, little attention has been paid to the distribution of human opinions. |
| Approach: | They use a dataset with 464,500 annotations to study Collective HumAn OpinionS . they argue that models lack the ability to recover the distribution over human labels . |
| Outcome: | The proposed dataset examines the distribution of human opinions in NLU evaluation datasets. |
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)
Copied to clipboard
Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta, Sarah Luger, Tharindu Ranasinghe, Ashiqur KhudaBukhsh, Marcos Zampieri, Christopher Homan
| Challenge: | Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values. |
| Approach: | They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data. |
| Outcome: | The proposed method breaks down disagreements by asking raters how they think others would annotate the data. |
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to assess annotator judgments aggregate disagreement into a single gold label . et al., 2022) show disagreement is diffuse, but standard approaches are not rigorous . |
| Approach: | They propose a schema-level diagnostic for auditing expert-designed annotation schemas prior to gold-label commitment . they find disagreement is not diffuse: instability concentrates in a few criteria, while nearly half of covered sentences activate multiple categories. |
| Outcome: | The proposed diagnostic separates unstable criteria with hard-to-operationalize boundaries and systematic overlap that blurs the boundaries between mutually exclusive categories. |