Challenge: Disagreement arises from subjective human opinion and can vary with one’s identity, beliefs, and social environment.
Approach: They propose a human-centered framework for reproducible ML evaluation and AI alignment that takes disagreement into account when building human-centric AI systems.
Outcome: The proposed framework is based on a human-centered and perspective-aware framework for reproducible ML evaluation and AI alignment.

Similar Papers

Towards Multi-Perspective NLP Systems: A Thesis Proposal (2025.acl-srw)

Copied to clipboard

Challenge: Existing approaches to resolving disagreements ignore individual opinions and can result in the marginalization of minority perspectives.
Approach: They propose to preserve individual labels in human-annotated datasets for subjective tasks and propose solutions for developing Perspective-Aware by design systems.
Outcome: The proposed framework will be used to develop more responsible and inclusive models.
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: In this position paper, we argue that human evaluation of generative large language models (LLMs) should be a multidisciplinary undertaking that draws upon the insights from disciplines such as user experience research and human behavioral psychology to ensure that the results are reliable.
Approach: They propose a framework for human evaluation of generative large language models that takes into account usability, aesthetics and cognitive biases.
Outcome: The proposed framework is based on the framework proposed by Deutsch and alnajjar . it is aimed at ensuring that human evaluation is accurate in the age of generative AI .
Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures.
Approach: They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope.
Outcome: The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches .
Proposal: From One-Fit-All to Perspective Aware Modeling (2025.acl-srw)

Copied to clipboard

Challenge: Variation in human annotation and human perspectives has drawn increasing attention in natural language processing research.
Approach: They propose to use annotation formats that better capture granularity and uncertainty of individual judgments and annotation modeling that leverages socio-demographic features to better represent and predict underrepresented or minority perspectives.
Outcome: The proposed tasks aim to advance natural language processing research towards more faithfully reflecting the diversity of human interpretation, enhancing both inclusiveness and fairness in language technologies.
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation.
Approach: They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation.
Outcome: The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics.
Seeing All Sides: Multi-Perspective In-Context Learning for Subjective NLP (2026.findings-eacl)

Copied to clipboard

Challenge: Modern language models excel at factual reasoning but struggle with value diversity, authors say . task-sensitive tasks such as hate speech expose this limitation . human disagreement captures the diversity of plausible human perspectives, authors argue .
Approach: They evaluate four large language models with human disagreements on five datasets . they find multi-perspective in-context learning outperforms standard prompting .
Outcome: The proposed approach outperforms standard prompting on English labels while disaggregated soft predictions better align with human judgments in Arabic and Italian datasets.
Aligning Language Models to User Opinions (2023.findings-emnlp)

Copied to clipboard

Challenge: Personality is a defining feature of human beings, shaped by a complex interplay of demographic characteristics, moral principles, and social experiences.
Approach: They use public opinion surveys to model past user opinions in addition to user demographics and ideology to achieve up to 7 points accuracy gains in predicting public opinions from survey questions.
Outcome: The proposed model achieves 7 points accuracy gains in predicting public opinions from public opinion surveys across a broad set of topics.
Aligning Black-box Language Models with Human Judgments (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks.
Approach: They propose a framework to align LLM judgments with individual human evaluators or their aggregated judgments without retraining or fine-tuning the LLM.
Outcome: The proposed framework achieves 142% improvement in agreement across 29 tasks and exceeds inter-human agreement on four out of six tasks.
Humans or LLMs as the Judge? A Study on Judgement Bias (2024.emnlp-main)

Copied to clipboard

Challenge: Proprietary models such as GPT-4, Claude, Gemini-Pro and others are being democratized to improve evaluations of LLMs.
Approach: They propose a framework that is free from referencing groundtruth annotations for investigating **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia's** on LLM and human judges.
Outcome: The proposed framework investigates **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia' on LLM and human judges.
To Mask or to Mirror: Human-AI Alignment in Collective Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used to model and augment collective decision-making.
Approach: They propose a framework for assessing collective alignment using the Lost at Sea social psychology task.
Outcome: The proposed framework compares LLMs with human-AI alignment on the Lost at Sea social psychology task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations