Challenge: Human annotator simulation (HAS) is a cost-effective alternative to human evaluation tasks.
Approach: They propose a framework to model human annotation variability via meta-learning . conditional softmax flow model leverages diverse human annotations via meta learning . results demonstrate that method can predict aggregated behaviours of human annotators .
Outcome: The proposed method achieves state-of-the-art performance on two real-world human evaluation tasks: emotion recognition and toxic speech detection.

Similar Papers

Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation (2024.findings-emnlp)

Copied to clipboard

Challenge: In subjective tasks, the inclusion of diverse annotators is crucial as their unique perspectives significantly influence the annotations.
Approach: They propose a framework that minimizes the annotation budget while maximizing the predictive performance for each annotator.
Outcome: The proposed framework surpasses the previous SOTA in capturing the annotators’ individual perspectives with as little as 25% of the original annotation budget on two datasets.
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation.
Approach: They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation.
Outcome: The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics.
HUMAN: Hierarchical Universal Modular ANnotator (2020.emnlp-demos)

Copied to clipboard

Challenge: HUMAN is a web-based annotation tool that covers a variety of annotation tasks on textual and image data.
Approach: They propose a web-based annotation tool that covers a variety of annotation tasks on textual and image data.
Outcome: HUMAN covers a variety of annotation tasks on textual and image data and uses an internal deterministic state machine to chain different tasks in an interdependent manner.
The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics (2023.acl-short)

Copied to clipboard

Challenge: Existing work has attempted to model individual annotation behaviour rather than predicting aggregated labels.
Approach: They propose to model individual annotator behaviour rather than predicting aggregated labels by adding group-specific layers to multi-annotator models to account for sociodemographics.
Outcome: The proposed model does not significantly improve on toxic content detection tasks.
Modeling Human Perspectives with Socio-Demographic Representations (2026.findings-acl)

Copied to clipboard

Challenge: Recent studies show that human disagreement is widespread across many annotation tasks.
Approach: They propose a method that jointly models annotator perspectives while learning socio-demographic representations.
Outcome: The proposed method outperforms concatenation-based methods in predicting annotator perspectives . it learns socio-demographic representations and analyzes how demographic factors relate to variation .
Label and Explanation Variation in LLM-Based Annotation: a Case Study in Natural Language Inference (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown considerable promise for annotation purposes, but questions remain about their ability to capture human label variation (HLV) label variation is genuine disagreement between annotators observed across NLP tasks.
Approach: They investigate how label and explanation variation manifests within and across LLMs with respect to the Natural Language Inference task.
Outcome: The proposed models generate label distributions similar to humans but exhibit distinct, idiosyncratic judgments and disagreement patterns.
Dynamic Human Evaluation for Relative Model Comparisons (2022.lrec-1)

Copied to clipboard

Challenge: Automated metrics have reported flaws when applied to measure quality aspects of generated text and have been shown to correlate poorly with human judgements.
Approach: They propose an agent-based framework to measure the required number of human annotations when evaluating generated outputs in relative comparison settings.
Outcome: The proposed model can be compared with a crowdsourced case study and a simulation with simulated human judgements.
The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: a paper argues that human label variation impacts all stages of the ML pipeline . human label variations are often considered noise due to disagreement, subjectivity in annotation or multiple plausible answers.
Approach: They propose to reconcile different notions of human label variation and propose a repository of publicly-available datasets with un-aggregated labels.
Outcome: The proposed approaches are compared with publicly available datasets with un-aggregated labels and identify gaps.
Unveiling the Multi-Annotation Process: Examining the Influence of Annotation Quantity and Instance Difficulty on Model Performance (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have shown that multi-annotator datasets can improve performance when they expand from a single annotation per instance to multiple annotations.
Approach: They propose a multi-annotator simulation process to generate datasets with varying annotation budgets and compare them to a single annotation per instance.
Outcome: The proposed model can generate datasets with varying annotation budgets and show that similar datasets can lead to varying performance gains.
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals’ Subjective Text Perceptions (2025.acl-long)

Copied to clipboard

Challenge: Recent work has shown that LLMs perform poorly when prompted with sociodemographic attributes, suggesting limited inherent sociodemography knowledge.
Approach: They propose to train large language models to be accurate sociodemographic models of annotator variation by using a curated dataset of five tasks with standardized sociodemography.
Outcome: The proposed models improve in sociodemographic prompting when trained but this performance gain is largely due to models learning annotator-specific behaviour rather than sociodemography.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations