Challenge: Proposed solutions aim to capture minority perspectives by either modelling annotator disagreements or grouping annotators based on shared metadata.
Approach: They propose a framework that trains models without encoding annotator metadata and creates clusters of similar opinions, that are called voices.
Outcome: The proposed framework captures minority perspectives based on demographic factors in two distinct datasets while also capturing majority perspectives.

Similar Papers

Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning (2023.acl-long)

Copied to clipboard

Challenge: Annotator disagreements are resolved before learning takes place, but researchers question the performance of a system when annotators disagree.
Approach: They propose a method that uses language features and label distributions to pool similar items into larger labels.
Outcome: The proposed method is based on five publicly available datasets with varying levels of disagreements on social media and in the wild using a dataset from Facebook.
Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets (D19-1)

Copied to clipboard

Challenge: Having only a few workers generate the majority of dataset examples raises concerns about data diversity .
Approach: They perform a series of experiments to investigate annotator biases in recent NLU datasets . they find that models are able to recognize the most productive annotators .
Outcome: The results show that models can recognize the most productive annotators and do not generalize well to examples from annotator that did not contribute to the training set.
Proposal: From One-Fit-All to Perspective Aware Modeling (2025.acl-srw)

Copied to clipboard

Challenge: Variation in human annotation and human perspectives has drawn increasing attention in natural language processing research.
Approach: They propose to use annotation formats that better capture granularity and uncertainty of individual judgments and annotation modeling that leverages socio-demographic features to better represent and predict underrepresented or minority perspectives.
Outcome: The proposed tasks aim to advance natural language processing research towards more faithfully reflecting the diversity of human interpretation, enhancing both inclusiveness and fairness in language technologies.
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)

Copied to clipboard

Challenge: The first workshop on crowdsourcing for NLP is open to all .
Approach: The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks.
Outcome: The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data .
You Are What You Annotate: Towards Better Models through Annotator Representations (2023.findings-emnlp)

Copied to clipboard

Challenge: Annotator disagreement is ubiquitous in natural language processing tasks.
Approach: They propose to model annotators' idiosyncrasies and account for their idioms by creating representations for each annotator and their annotations.
Outcome: The proposed model improves on an existing dataset with eight annotators with inherent disagreements while increasing model size by 1%.
Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to label aggregation fail to capture subjective annotations and can lead to biases.
Approach: They propose annotator-aware representations for text for subjective classification tasks that involve learning representations of annotators.
Outcome: The proposed model improves on metrics that assess the performance on capturing individual annotators’ perspectives.
Learning from Measurements in Crowdsourcing Models: Inferring Ground Truth from Diverse Annotation Types (C18-1)

Copied to clipboard

Challenge: Annotated corpora are often assigned to internet workers whose judgments are reconciled by crowdsourcing models.
Approach: They propose a framework for learning from rich prior knowledge to combine annotations with different structures.
Outcome: The proposed model compares favorably with previous work and enables active sample selection to reduce annotation effort.
Modeling Human Perspectives with Socio-Demographic Representations (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies show that human disagreement is widespread across many annotation tasks.
Approach: They propose a method that jointly models annotator perspectives while learning socio-demographic representations.
Outcome: The proposed method outperforms concatenation-based methods in predicting annotator perspectives . it learns socio-demographic representations and analyzes how demographic factors relate to variation .
Corpus Considerations for Annotator Modeling and Scaling (2024.naacl-long)

Copied to clipboard

Challenge: Recent trends in natural language processing and annotation tasks emphasize individual perspectives . annotator models that rely on a single ground truth may disregard valuable minority perspectives omissions .
Approach: They propose a composite embedding approach to investigate annotator modeling techniques . they show that the commonly used user token model consistently outperforms more complex models .
Outcome: The proposed model outperforms more complex models on a given dataset.
A Streamlined Method for Sourcing Discourse-level Argumentation Annotations from the Crowd (N19-1)

Copied to clipboard

Challenge: Existing methods for analyzing discourse-level argument annotations require expensive labor and data.
Approach: They propose a method that breaks down a popular but complex discourse-level argument annotation scheme into a simple iterative procedure that can be applied even by untrained annotators.
Outcome: The proposed method can be applied even by untrained annotators.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations