Arbiters of Ambivalence: Challenges of using LLMs in No-Consensus tasks (2025.findings-acl)
Copied to clipboard
| Challenge: | LLMs are increasingly being used to replace humans in "aligning" LLM training . studies question this trend, but have found they can be more effective in ambivalent scenarios where humans disagree . |
| Approach: | They develop a “no-consensus” benchmark by curating examples that encompass a variety of a priori ambivalent scenarios. |
| Outcome: | The proposed benchmarks show that LLMs can provide nuanced assessments when generating open-ended answers, but tend to take a stance on no-consensus topics when employed as judges or debaters. |
Similar Papers
Wait, that’s not an option: LLMs Robustness with Incorrect Multiple-Choice Options (2025.acl-long)
Copied to clipboard
| Challenge: | Using a framework that combines instruction-following with critical reasoning, we show that the ability of LLMs to override defaults when faced with invalid options is impaired by alignment techniques. |
| Approach: | They propose a framework for evaluating LLMs’ capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers. |
| Outcome: | The proposed framework improves models' ability to override defaults when faced with invalid options while minimizing the impact of model size and training techniques on the model. |
Dissecting Human and LLM Preferences (2024.acl-long)
Copied to clipboard
| Challenge: | a recent study shows that human and Large Language Model preferences are important for model fine-tuning and evaluation. |
| Approach: | They dissect the preferences of human and 32 different Large Language Models to understand their quantitative composition. |
| Outcome: | The proposed model is compared with 32 different large language models using real-world user-model conversations. |
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)
Copied to clipboard
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessandro Suglia, Aditya K Surikuchi, Ece Takmaz, Alberto Testoni
| Challenge: | Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models . |
| Approach: | They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets. |
| Outcome: | The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets. |
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference (2025.findings-naacl)
Copied to clipboard
| Challenge: | Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. |
| Approach: | They propose a system-level evaluation framework that ranks LLMs based on their alignment with human preferences. |
| Outcome: | The proposed framework aims to rank LLMs based on their performance and alignment with human preferences. |
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly studied as repositories of linguistic knowledge. |
| Approach: | They compare LLMs’ performance as pragmatic listeners and as pragmatic speakers . they find a robust asymmetry between pragmatic evaluation and pragmatic generation . |
| Outcome: | The proposed models perform better as listeners than speakers, and produce more appropriate language than speakers. |
Humans or LLMs as the Judge? A Study on Judgement Bias (2024.emnlp-main)
Copied to clipboard
| Challenge: | Proprietary models such as GPT-4, Claude, Gemini-Pro and others are being democratized to improve evaluations of LLMs. |
| Approach: | They propose a framework that is free from referencing groundtruth annotations for investigating **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia's** on LLM and human judges. |
| Outcome: | The proposed framework investigates **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia' on LLM and human judges. |
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies. |
| Approach: | They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
| Outcome: | The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judge (2025.findings-emnlp)
Copied to clipboard
Chiyu Ma, Enpei Zhang, Yilun Zhao, Wenjun Liu, Yaning Jia, Peijun Qing, Lin Shi, Arman Cohan, Yujun Yan, Soroush Vosoughi
| Challenge: | LLM-as-Judge frameworks provide scalable alternative to human evaluation . but the question of how intrinsic biases manifest in these settings remains unexplored . |
| Approach: | They conduct systematic analysis of four bias types in multi-agent LLM-as-Judge frameworks . they find debate framework amplifies biases sharply after initial debate . |
| Outcome: | The proposed frameworks amplify biases after debate and show they are stronger in meta-judge scenarios. |
DebateQA: Evaluating Question Answering on Debatable Knowledge (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing QA benchmarks that provide fixed answers to debatable questions are inadequate for evaluating their performance. |
| Approach: | They propose to use a dataset of 2,941 debatable questions to assess their ability to provide comprehensive answers to inherently debatably asked questions. |
| Outcome: | The proposed model performs well on 2,941 debatable questions accompanied by human-annotated partial answers that capture a variety of perspectives. |
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios . |
| Approach: | They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs . |
| Outcome: | The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB . |