| Challenge: | a wide consensus is rife regarding the need for reference annotated datasets . however, the creation of such datasets is accompanied by theorectical and practical issues . |
| Approach: | They propose to use agreement among annotators as an indicator of consensus . they argue that it is difficult to produce gold-standard annotated datasets . |
| Outcome: | The proposed model focuses on the complex relations between agreement and reference and the emergence of consensus. |
Similar Papers
Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement (2020.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation methods rely on agreement between annotators, which implies a single correct interpretation. |
| Approach: | They propose an agreement-independent quality metric based on answer-coherence to evaluate on expected disagreement. |
| Outcome: | The proposed model shows that agreement is the most important indicator of quality in semantic annotation tasks. |
Verifying Annotation Agreement without Multiple Experts: A Case Study with Gujarati SNACS (2023.findings-acl)
Copied to clipboard
| Challenge: | a small fraction of the about 7,000 languages of the world have datasets or linguistic tools . linguistic datasets are a foundation of NLP research, but they are not always reliable . authors propose weak verifiers to help estimate dataset quality . |
| Approach: | They propose four weak verifiers to help estimate dataset quality . they propose to use Gujarati as a low-resource language to test for dataset quality. |
| Outcome: | The proposed methods concur with a double-annotation study in Gujarati. |
Common Law Annotations: Investigating the Stability of Dialog System Output Annotations (2023.findings-acl)
Copied to clipboard
Seunggun Lee, Alexandra DeLucia, Nikita Nangia, Praneeth Ganedi, Ryan Guan, Rubing Li, Britney Ngaw, Aditya Singhal, Shalaka Vaidya, Zijun Yuan, Lining Zhang, João Sedoc
| Challenge: | High agreement is often used to show reliability of annotation procedures, but it is insufficient to ensure or reproducibility. |
| Approach: | They propose a protocol that increases Inter-Annotator Agreement among annotators and a standardized and codified protocol that strictly enforces transparency in the annotation process. |
| Outcome: | The proposed protocol ensures transparency in the annotation process, which ensures reproducibility of annotation guidelines. |
Establishing Annotation Quality in Multi-label Annotations (2022.coling-1)
Copied to clipboard
| Challenge: | Multi-label annotations allow multiple interpretations of a single item, but they also affect the chance that two coders agree with each other. |
| Approach: | They propose a bootstrapped method to obtain chance agreement for each measure and a method to get an adjusted agreement coefficient that is more interpretable. |
| Outcome: | The proposed method allows for an adjusted agreement coefficient that is more interpretable on simulated datasets. |
Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)
Copied to clipboard
| Challenge: | Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer. |
| Approach: | They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines. |
| Outcome: | The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures. |
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)
Copied to clipboard
| Challenge: | a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes . |
| Approach: | They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones . |
| Outcome: | The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions. |
Let’s discuss! Quality Dimensions and Annotated Datasets for Computational Argument Quality Assessment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Argumentation is a key competence and an important cultural technique in democratic societies. |
| Approach: | They propose to create domain-specific datasets and methods to assess argument quality. |
| Outcome: | The proposed methods address gaps in the literature and aid future research in the domain. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
Rethinking the Agreement in Human Evaluation Tasks (C18-1)
Copied to clipboard
| Challenge: | In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability. |
| Approach: | They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias. |
| Outcome: | The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks. |