Challenge: Disagreement in natural language annotation has been studied from a perspective of biases introduced by the annotators and the annotation frameworks.
Approach: They propose to analyze task design bias in crowdsourced annotations where lay annotators are used to elicit interpretations.
Outcome: The proposed methods can push annotators towards certain relations and some discourse relation senses can be better elicited with one or the other approach.

Similar Papers

Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets (D19-1)

Copied to clipboard

Challenge: Having only a few workers generate the majority of dataset examples raises concerns about data diversity .
Approach: They perform a series of experiments to investigate annotator biases in recent NLU datasets . they find that models are able to recognize the most productive annotators .
Outcome: The results show that models can recognize the most productive annotators and do not generalize well to examples from annotator that did not contribute to the training set.
Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training (2022.lrec-1)

Copied to clipboard

Challenge: Recent methods have obtained promising results by extracting relation labels from participants . obtaining linguistic annotations from novice crowdworkers is difficult . crowdsourcing allows for fast and cost-effective collection of labelled data, but because tasks need to be intuitive, crowdworker cannot be asked to perform them.
Approach: They propose to use a selection-only approach to obtain linguistic annotations from novices . current study shows that the method is cost- and time-intensive .
Outcome: The current study shows that selection and training improves the agreement between workers and gold labels, but the method is cost- and time-intensive.
Mining Crowdsourcing Problems from Discussion Forums of Workers (2020.coling-main)

Copied to clipboard

Challenge: Among the most widely used platforms are Upwork, Appen, and above all Amazon Mechanical Turk (MTurk) which host annotation tasks and collect huge sets of annotated data from workers.
Approach: They propose to use topic modeling to analyze workers' complaints from a new English corpus of workers’ forum discussions to identify problems in task design, task operation, and task evaluation that workers face with requesters in crowdsourcing processes.
Outcome: The findings form the basis for future research on how to improve crowdsourcing processes.
DiscoGeM: A Crowdsourced Corpus of Genre-Mixed Implicit Discourse Relations (2022.lrec-1)

Copied to clipboard

Challenge: DiscoGeM is a crowdsourced corpus of 6,505 implicit discourse relations . the results show that a significant proportion of discourse relations are ambiguous . text genre is crucially affected by the distribution of discourse relation labels .
Approach: They propose to use crowdsourced corpus of 6,505 implicit discourse relations to classify relations . they propose to include genre as a factor in automatic relation classification .
Outcome: The proposed dataset shows that a significant proportion of discourse relations are ambiguous and can express multiple relation senses.
Toward Annotator Group Bias in Crowdsourcing (2022.acl-long)

Copied to clipboard

Challenge: Annotator group bias is a common problem in crowdsourcing, but is often overlooked .
Approach: They propose a probabilistic framework to capture annotator group bias using an extended Expectation Maximization algorithm.
Outcome: The proposed model can model annotator group bias over competitive datasets and demonstrate that it is effective over multiple datasets.
Context Does Matter: Implications for Crowdsourced Evaluation Labels in Task-Oriented Dialogue Systems (2024.findings-naacl)

Copied to clipboard

Challenge: Existing studies suggest using only a portion of the dialogue context in the annotation process, but the impact of this limitation on label quality remains unexplored.
Approach: They propose to use large language models to summarize the dialogue context to provide a rich and short description of the dialogue and to examine the impact of doing so on the annotator’s performance.
Outcome: The proposed model reduces the context and produces higher quality ratings but introduces ambiguity in usefulness ratings.
Improving Crowdsourcing-Based Annotation of Japanese Discourse Relations (L18-1)

Copied to clipboard

Challenge: Discourse parsing is an important task in natural language processing, but few languages have corpora annotated with discourse relations . crowdsourcing-based annotations are of poor quality and require expensive and time-consuming . et al. (2009) evaluated the quality of annotations using expert annotations.
Approach: They construct a Japanese corpus with discourse annotations through crowdsourcing . they propose improvement techniques based on language tests .
Outcome: The proposed methods improve the quality of the annotations, and will make them publicly available.
Task Assignment meets Annotator Modeling: Human-LLM Collaborative Annotation with Constraints (2026.acl-srw)

Copied to clipboard

Challenge: Existing approaches to label annotation are labor-intensive and time-consuming.
Approach: They propose a framework that estimates per-task accuracy from task features using a learning from crowds model and incorporates these estimations into a linear programming formulation that assigns tasks under practical constraints.
Outcome: The proposed method achieves comparable accuracy to baseline methods while satisfying given constraints.
Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference (2025.coling-main)

Copied to clipboard

Challenge: Creating NLP datasets with Large Language Models (LLMs) is an attractive alternative to relying on crowd-source workers.
Approach: They recreate a portion of the Stanford Natural Language Inference corpus using GPT-4, Llama-2 70b for Chat, and Mistral 7b Instruct.
Outcome: The proposed model can be used to generate NLP datasets with stereotypical biases and annotation artifacts.
Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks (2023.eacl-main)

Copied to clipboard

Challenge: Disagreement can reflect different aspects of linguistic annotation, from annotators’ subjectivity to sloppiness or lack of context to interpret a text.
Approach: They propose a taxonomy of possible reasons leading to annotators' disagreement in subjective tasks and manually label part of a Twitter dataset for offensive language detection in english following this taxonomies.
Outcome: The proposed taxonomy of disagreements in linguistic datasets can be used to assess how accurate tweets belonging to different disagreement categories can be classified as offensive or not.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations