Challenge: Large Language Models (LLMs) are increasingly embedded in the daily lives of individuals across diverse social classes.
Approach: They propose to analyze LLMs' responses to 1,016 scenarios categorized into ethical, unethical, and neutral types.
Outcome: The proposed model analyzed 1,016 scenarios categorized into ethical, unethical, and neutral types.

Similar Papers

The discordance between embedded ethics and cultural inference in large language models (2025.emnlp-main)

Copied to clipboard

Challenge: Effective interactions between AI and humans require an accurate representation of diverse cultures.
Approach: They propose a framework that embeds ethical principles within an LLM and a hyperplane that embedding cultural norms within it.
Outcome: The proposed framework shows that cultural norms are more aligned with ethical principles than standard models.
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception (2025.coling-main)

Copied to clipboard

Challenge: Detecting media bias is critical due to the spread of misinformation and disinformation on social media platforms.
Approach: They investigate the presence and nature of bias within large language models and its consequential impact on media bias detection.
Outcome: The proposed debiasing strategies include prompt engineering and model fine-tuning.
Neutral Is Not Unbiased: Evaluating Implicit and Intersectional Identity Bias in LLMs Through Structured Narrative Scenarios (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models often reproduce societal biases, yet most evaluations overlook how such biase evolve across nuanced contexts or intersecting identities.
Approach: They propose a scenario-based evaluation framework built on 100 narrative tasks . they use critical discourse analysis and quantitative linguistic metrics to analyze LLMs .
Outcome: The proposed evaluation framework provides ethically coherent and socially plausible settings for probing model behavior.
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Representative bias is a tendency of Large Language Models to generate outputs that mirror the experiences of certain identity groups, and affinity bias is an evaluative preference for specific narratives.
Approach: They propose two new metrics to measure representative bias and affinity bias within large language models and present a new set of tasks designed with customized rubrics to detect these biases.
Outcome: The proposed model identifies representative biases in prominent LLMs, with a preference for identities associated with being white, straight, and men.
The Impossibility of Fair LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing frameworks for evaluating large language models do not extend to general-purpose AI contexts or are infeasible in practice.
Approach: They analyze a variety of technical fairness frameworks to find inherent challenges . they find that each framework does not logically extend to the general-purpose AI context .
Outcome: The proposed frameworks do not logically extend to the general-purpose AI context or are infeasible in practice due to large amounts of unstructured training data and potential combinations of human populations, use cases, and sensitive attributes.
Addressing Bias and Hallucination in Large Language Models (2024.lrec-tutorials)

Copied to clipboard

Challenge: This tutorial provides a comprehensive overview of two critical aspects of Large Language Models: bias and hallucination.
Approach: This tutorial provides an overview of two critical aspects of Large Language Models: bias and hallucination.
Outcome: This tutorial delves into the complex dimensions of Large Language Models (LLMs) it outlines ethical considerations pertinent to their development and discusses hallucination, a prevalent issue in generative AI systems such as LLMs.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Unmasking Implicit Bias: Evaluating Persona-Prompted LLM Responses in Power-Disparate Social Scenarios (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in simulating human behaviour and social intelligence, but they risk perpetuating societal biases, especially when demographic information is involved.
Approach: They propose a framework that measures semantic shifts in responses and an LLM-judged Preference Win Rate to assess how demographic prompts affect response quality across power-disparate social scenarios.
Outcome: The proposed framework measures semantic shifts in responses and an LLM-judged Preference Win Rate (WR) to assess how demographic prompts affect response quality across power-disparate social scenarios.
Investigating Subtler Biases in LLMs: Ageism, Beauty, Institutional, and Nationality Bias in Generative Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in language generation models can be used to assist users in a variety of tasks, but there are risks associated with introducing LLM biases into consequential decisions.
Approach: They propose to use a template-generated dataset to measure subtler correlated decisions that LLMs make between social groups and unrelated positive and negative attributes.
Outcome: The proposed model can be used to evaluate progress in more generalized biases and extend the benchmark with minimal human annotation.
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say .
Approach: They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.
Outcome: The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations