Challenge: Large language models (LLMs) can exhibit political biases, which creates a risk of undue influence on LLM users and public opinion.
Approach: They use a dataset of 36k real-time test prompts to measure LLM political bias on U.S. and Chinese issues.
Outcome: The proposed model origin and prompt language influence bias on 60 political issues.

Similar Papers

Navigating the Political Compass: Evaluating Multilingual LLMs across Languages and Nationalities (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are ubiquitous in today’s technological landscape, boasting a plethora of applications, and even endangering human jobs in complex and creative fields.
Approach: They evaluate the political bias of 15 multilingual LLMs using the Political Compass Test and assign a nationality to each model.
Outcome: The models on the 50 most populous countries and their official languages exhibit political bias.
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and measures focus on gender and racial biases, but political bias exists in LLMs and can lead to polarization and other harms in downstream applications.
Approach: They propose to analyze the content and style of LLMs generated by political issues and propose a framework that can be scalable to other topics.
Outcome: The proposed framework is easily scalable to other topics and is explainable.
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Large language models exhibit cultural and geopolitical biases when their outputs shape public opinion or reinforce dominant narratives.
Approach: They define two types of bias in large language models: model bias and inference bias through a two-phase evaluation.
Outcome: The proposed framework evaluates large language models on factual and disputable questions across four languages and question types.
Biased LLMs can Influence Political Decision-Making (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have found that biased LLMs can influence decisions in areas such as medical classifications and educational hiring.
Approach: They conducted two interactive experiments on partisan bias in large language models while completing tasks with either a biased liberal, biased conservative, or unbiased control model.
Outcome: The results show that prior knowledge of AI is weakly correlated with a reduction of the bias, suggesting that AI education can be crucial for mitigating bias effects.
Measuring and Mitigating Media Outlet Name Bias in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have explored the potential political biases of large language models, but limited attention has been devoted to the effects of media outlet names.
Approach: They propose to quantify media outlet name biases in large language models and leverage this metric to develop an automated prompt optimization framework.
Outcome: The proposed framework mitigates media outlet name biases, offering a scalable approach to enhancing the fairness of LLMs in news-related applications.
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to analyze political biases rely on small-size intermediate tasks and the LLMs themselves.
Approach: They propose an entropy-based inconsistency metric to encode political biases . they insert 1319 demographically and politically diverse politician names in 450 political sentences .
Outcome: The proposed method combines high accuracy with a correct understanding of the candidate candidate.
Investigating Bias in LLM-Based Bias Detection: Disparities between LLMs and Human Perception (2025.coling-main)

Copied to clipboard

Challenge: Detecting media bias is critical due to the spread of misinformation and disinformation on social media platforms.
Approach: They investigate the presence and nature of bias within large language models and its consequential impact on media bias detection.
Outcome: The proposed debiasing strategies include prompt engineering and model fine-tuning.
Assessing Reliability and Political Bias In LLMs’ Judgements of Formal and Material Inferences With Partisan Conclusions (2025.acl-long)

Copied to clipboard

Challenge: This paper examines the ability of LLMs to correctly label simple inferences with partisan conclusions.
Approach: They develop a dataset with formal and material inferences with conclusions that favor either the political left or the political right.
Outcome: The proposed models show that they are unreliable and political bias persists throughout the English and German datasets.
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance (2026.tacl-1)

Copied to clipboard

Challenge: Large language models are helping millions of users write texts about diverse issues . issue bias is where an LLM tends to present just one perspective on a given issue .
Approach: They construct a set of 2.49m realistic English-language prompts to measure issue bias in LLM writing assistance using 3.9k templates and 212 political issues from real user interactions.
Outcome: The proposed model aligns more with US Democrat than Republican voter opinion on a subset of issues.
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks (2025.acl-long)

Copied to clipboard

Challenge: Existing work on large language models lacks robustness, highlighting the limitations of such models.
Approach: They propose a novel approach where two LLMs engage in self-debate to persuade a neutral version of the model.
Outcome: The proposed approach examines whether large language models are robust during interactions and whether they are susceptible to reinforcing misinformation or shifting to harmful viewpoints.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations