Challenge: Proprietary models such as GPT-4, Claude, Gemini-Pro and others are being democratized to improve evaluations of LLMs.
Approach: They propose a framework that is free from referencing groundtruth annotations for investigating **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia's** on LLM and human judges.
Outcome: The proposed framework investigates **Misinformation Oversight Bias**, **Gender Bia**,**Authority Bia* and **Beauty Bia' on LLM and human judges.

Similar Papers

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (2025.acl-short)

Copied to clipboard

Challenge: Existing evaluations of NLP models with LLMs are based on human judgments . however, there are concerns about their validity and reproducibility in proprietary models .
Approach: They evaluate 11 current LLMs for their ability to replicate annotations. they show substantial variance across models and datasets.
Outcome: The proposed model can replicate human annotations on 20 NLP datasets and show substantial variance across models and datasets.
Don’t Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as evaluators for code evaluation tasks . however, whether they can handle superficial variations remains unclear .
Approach: They define six types of potential biases in code evaluation and reveal their impact on LLM judges.
Outcome: The proposed method can be used to evaluate semantically equivalent code with superficial variations without reference implementations.
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios .
Approach: They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios.
Outcome: The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm.
How Reliable is Multilingual LLM-as-a-Judge? (2025.findings-emnlp)

Copied to clipboard

Challenge: LLMs are a popular evaluation strategy, but their reliability in multilingual evaluation remains uncertain.
Approach: They evaluate five models from different model families across five diverse tasks involving 25 languages.
Outcome: The models perform poorly across languages and average Fleiss’ Kappa is 0.3 .
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks (2025.acl-long)

Copied to clipboard

Challenge: Existing work on large language models lacks robustness, highlighting the limitations of such models.
Approach: They propose a novel approach where two LLMs engage in self-debate to persuade a neutral version of the model.
Outcome: The proposed approach examines whether large language models are robust during interactions and whether they are susceptible to reinforcing misinformation or shifting to harmful viewpoints.
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models carry implicit biases across race, gender, and religion . prior studies documented such biase based on text generation and classification tasks .
Approach: They investigate bias in large language models by controlling metadata on author metadata . authors found affiliation bias favoring authors from highly ranked institutions .
Outcome: The proposed model favors authors from highly ranked institutions, the authors show . the model also favors author affiliations from highly-ranked institutions .
JuStRank: Benchmarking LLM Judges for System Ranking (2025.acl-long)

Copied to clipboard

Challenge: Recent work has focused on instance-based evaluation of LLM judges, where a judge is evaluated over a set of responses, or response pairs, while being agnostic to their source systems.
Approach: They propose to validate the quality of the LLM judge itself by comparing system scores to a human-based ranking.
Outcome: The proposed model fails to validate the quality of the judge itself, ignoring critical factors affecting system-level ranking, such as a judge’s positive or negative bias towards certain systems.
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception (2025.coling-main)

Copied to clipboard

Challenge: Detecting media bias is critical due to the spread of misinformation and disinformation on social media platforms.
Approach: They investigate the presence and nature of bias within large language models and its consequential impact on media bias detection.
Outcome: The proposed debiasing strategies include prompt engineering and model fine-tuning.
Can Large Language Models Be an Alternative to Human Evaluations? (2023.acl-long)

Copied to clipboard

Challenge: Human evaluation is indispensable for assessing the quality of texts generated by machine learning models or written by humans.
Approach: They propose to use large language models to evaluate unseen texts using the same instructions and samples . they also use LLMs to generate responses to questions that are used to conduct human evaluation .
Outcome: The proposed model can be used to evaluate texts in open-ended story generation and adversarial attacks.
Can LLM be a Personalized Judge? (2024.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the reliability of large language models (LLMs) for personalization and role-playing evaluation without examining its validity.
Approach: They investigate the reliability of LLM-as-a-Personalized-Judge for personalization . they find that personas provided to LLMs have limited predictive power .
Outcome: The proposed model is less reliable than previously thought, the authors show . human annotation reveals that third-person crowd worker evaluations of personalized preferences are even worse than LLM predictions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations