Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)
Copied to clipboard
| Challenge: | Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer. |
| Approach: | They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines. |
| Outcome: | The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures. |
Similar Papers
Automatic Argument Quality Assessment - New Datasets and Methods (D19-1)
Copied to clipboard
Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, Noam Slonim
| Challenge: | 6.3k arguments were collected from contributors of various levels, and are released as part of this work. |
| Approach: | They propose to use a language model to annotate arguments for argument ranking and argument-pair classification. |
| Outcome: | The proposed methods outperform state-of-the-art methods in the argument ranking task and argument-pair classification task. |
A Streamlined Method for Sourcing Discourse-level Argumentation Annotations from the Crowd (N19-1)
Copied to clipboard
| Challenge: | Existing methods for analyzing discourse-level argument annotations require expensive labor and data. |
| Approach: | They propose a method that breaks down a popular but complex discourse-level argument annotation scheme into a simple iterative procedure that can be applied even by untrained annotators. |
| Outcome: | The proposed method can be applied even by untrained annotators. |
Rethinking the Agreement in Human Evaluation Tasks (C18-1)
Copied to clipboard
| Challenge: | In natural language processing, IAA is often viewed as a means of assessing the quality of data on a task, in particular, the reliability. |
| Approach: | They propose a new approach to use agreement metrics in natural language generation evaluation tasks to reduce subjective bias. |
| Outcome: | The proposed approach is based on the inter-annotator agreement (IAA) of natural language generation tasks. |
Common Law Annotations: Investigating the Stability of Dialog System Output Annotations (2023.findings-acl)
Copied to clipboard
Seunggun Lee, Alexandra DeLucia, Nikita Nangia, Praneeth Ganedi, Ryan Guan, Rubing Li, Britney Ngaw, Aditya Singhal, Shalaka Vaidya, Zijun Yuan, Lining Zhang, João Sedoc
| Challenge: | High agreement is often used to show reliability of annotation procedures, but it is insufficient to ensure or reproducibility. |
| Approach: | They propose a protocol that increases Inter-Annotator Agreement among annotators and a standardized and codified protocol that strictly enforces transparency in the annotation process. |
| Outcome: | The proposed protocol ensures transparency in the annotation process, which ensures reproducibility of annotation guidelines. |
Annotating Arguments in a Corpus of Opinion Articles (2022.lrec-1)
Copied to clipboard
Gil Rocha, Luís Trigo, Henrique Lopes Cardoso, Rui Sousa-Silva, Paula Carvalho, Bruno Martins, Miguel Won
| Challenge: | Argument annotation is the process of exposing and justifying one's points of view, with the aim of conveying a logical reasoning through a set of semantically related propositions. |
| Approach: | They propose to use argumentative discourse units to annotate arguments in Portuguese using a multi-layered process to analyze the annotations produced. |
| Outcome: | The proposed model exploits the best practices identified in previous studies while fostering the potential use of the resulting annotated corpus for new purposes. |
Would you describe a leopard as yellow? Evaluating crowd-annotations with justified and informative disagreement (2020.coling-main)
Copied to clipboard
| Challenge: | Existing evaluation methods rely on agreement between annotators, which implies a single correct interpretation. |
| Approach: | They propose an agreement-independent quality metric based on answer-coherence to evaluate on expected disagreement. |
| Outcome: | The proposed model shows that agreement is the most important indicator of quality in semantic annotation tasks. |
Interannotator Agreement for Lexico-Semantic Annotation of a Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a method for lexico-semantic annotation of the Basic Corpus of Polish Metaphors is described . the procedure is composed of three steps: deciding whether a particular occurrence of a word is asemantics or strictly grammatical. |
| Approach: | They propose a procedure for lexico-semantic annotation of the Basic Corpus of Polish Metaphor . procedure corrects morphosyntactic annotation of part of corpus that is automatically annotated . |
| Outcome: | The proposed procedure corrects the morphosyntactic annotation of part of the corpus . it is composed of three steps: deciding whether a word is asemantic or strictly grammatical . preliminary results show that the procedure is adequate for the task . |
Machine-Aided Annotation for Fine-Grained Proposition Types in Argumentation (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 2016 debates and commentary contains 4,648 argumentative propositions annotated with fine-grained proposition types. |
| Approach: | They propose a machine learning-human workflow for annotating for four complex proposition types . they demonstrate with preliminary analysis of rhetorical strategies and structure in presidential debates . |
| Outcome: | The proposed method can be used by technical researchers seeking more nuanced representations of argument . it can also be used to analyze rhetorical strategies and structure in presidential debates . |
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement (2021.emnlp-main)
Copied to clipboard
| Challenge: | supervised learning is a key component of offensive language detection, but there is little attention given to the quality of annotated data. |
| Approach: | They propose to examine the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. |
| Outcome: | The proposed datasets show that annotators' agreement has a strong effect on classifiers performance and robustness. |
Argument Mining as a Text-to-Text Generation Task (2024.eacl-long)
Copied to clipboard
| Challenge: | Argument Mining (AM) aims to uncover the argumentative structures within a text. |
| Approach: | They propose a method that generates argumentatively annotated text using a pretrained encoder-decoder language model and a pre-trained decoder. |
| Outcome: | The proposed method achieves state-of-the-art performance on three types of benchmark datasets. |