Investigating Reasons for Disagreement in Natural Language Inference (2022.tacl-1)

Copied to clipboard

Challenge: Several disagreements in natural language inference (NLI) annotation are due to uncertainty in the sentence meaning, others to annotator biases and task artifacts.
Approach: They propose a 4-way classification approach and a multilabel classification approach for detecting disagreements in natural language inference annotations.
Outcome: The proposed model is more expressive and gives better recall of possible interpretations in the data.

Similar Papers

Identifying inherent disagreement in natural language inference (2021.naacl-main)

Copied to clipboard

Challenge: Natural language inference is the task of determining whether text is entailed, contradicted or unrelated to another piece of text.
Approach: They propose to tease systematic inferences from disagreement items by capturing modes in annotations to simulate uncertainty in the annotation process.
Outcome: The proposed approach performs statistically better than baselines on the CommitmentBank corpus in English.
Ecologically Valid Explanations for Label Variation in NLI (2023.findings-emnlp)

Copied to clipboard

Challenge: Human label variation exists in many natural language processing tasks, including NLI .
Approach: They build an English dataset of 1,415 ecologically valid explanations for 122 MNLI items . they find that people can systematically vary on their interpretation .
Outcome: The proposed dataset contains 1,415 ecologically valid explanations for 122 items . the results show that people can vary on interpretation and highlight differences .
Capture Human Disagreement Distributions by Calibrated Networks for Natural Language Inference (2022.findings-acl)

Copied to clipboard

Challenge: Previously, it's common to disregard it as noise or as a sign of poor-quality data, as their annotations are heavily based on personal experience and opinions.
Approach: They propose to capture the human disagreement distribution from the perspective of model calibration.
Outcome: The proposed model can achieve competitive performance when well-calibrated, on divergence scores between predictive probability and the true human opinion distribution, and the accuracy.
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations (2026.findings-acl)

Copied to clipboard

Challenge: Natural Language Inference (NLI) datasets often exhibit label variation.
Approach: They extend LiTEx taxonomy to two NLI datasets and jointly analyze label variation and label variation.
Outcome: The proposed model combines explanations as a lens to analyze variation in NLI annotations and examine individual differences in reasoning.
Conflicts in Texts: Data, Implications and Challenges (2025.findings-emnlp)

Copied to clipboard

Challenge: Conflicts in data could reflect complexity of situations, changes that need to be explained and dealt with, difficulties in data annotation, and mistakes in generated outputs.
Approach: This survey categorizes conflicting information into three key areas . they identify the areas where conflicting data can be ignored and undermine models' reliability and trustworthiness.
Outcome: The findings highlight key challenges and future directions for developing conflict-aware NLP systems that can reason over and reconcile conflicting information more effectively.
Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks (2023.eacl-main)

Copied to clipboard

Challenge: Disagreement can reflect different aspects of linguistic annotation, from annotators’ subjectivity to sloppiness or lack of context to interpret a text.
Approach: They propose a taxonomy of possible reasons leading to annotators' disagreement in subjective tasks and manually label part of a Twitter dataset for offensive language detection in english following this taxonomies.
Outcome: The proposed taxonomy of disagreements in linguistic datasets can be used to assess how accurate tweets belonging to different disagreement categories can be classified as offensive or not.
Annotation Artifacts in Natural Language Inference Data (N18-2)

Copied to clipboard

Challenge: Large-scale datasets for natural language inference are created by crowdsourcing annotations . authors show that success of natural language models to date has been overestimated .
Approach: They propose a method for crowdsourcing annotations to generate 3 new sentences based on a sentence (premise) they show that a simple text categorization model can correctly classify the hypothesis alone in about 67% of SNLI and 53% of MultiNLI .
Outcome: The proposed model can classify the hypothesis alone in 67% of SNLI and 53% of MultiNLI datasets.
Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations (2022.tacl-1)

Copied to clipboard

Challenge: Annotators may systematically disagree with one another, reflecting their individual biases and values, especially in the case of subjective tasks such as detecting affect, aggression, and hate speech.
Approach: They propose to combine multi-annotator models with multi-task based approaches to resolve disagreements between annotations and derive single ground truth labels.
Outcome: The proposed model outperforms majority voting and averaging methods and estimates uncertainty in predictions.
Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown considerable promise for annotation purposes, but questions remain about their ability to capture human label variation (HLV) label variation is genuine disagreement between annotators observed across NLP tasks.
Approach: They investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference task.
Outcome: The proposed models generate label distributions similar to humans but exhibit distinct, idiosyncratic judgments and disagreement patterns.
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on evaluating large language models' ability to handle disagreement cases.
Approach: They evaluate the performance of large language models in detecting offensive language at varying levels of agreement.
Outcome: The proposed model improves detection accuracy and model alignment with human judgment by using disagreement samples in training.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations