Challenge: Existing probing methods for evaluating LLMs with political questions have limited stability and are unreliable.
Approach: They propose to use human survey data as in-context examples to query LLMs with political questions to evaluate their potential biases.
Outcome: The proposed task improves the stability of question-based bias evaluation and may be used to compare instruction-tuned models to their base versions.

Similar Papers

Quantifying the Influence of Irrelevant Contexts on Political Opinions Produced by LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Recent studies have examined the generation of large language models (LLMs) on subjective topics such as political opinions and attitudinal questionnaires.
Approach: They use a Political Compass Test questionnaire to quantify how irrelevant information can systematically bias model opinions in specific directions.
Outcome: The results show that even seemingly unrelated contexts alter model responses in predictable ways.
A Dual-Layered Evaluation of Geopolitical and Cultural Bias in LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Large language models exhibit cultural and geopolitical biases when their outputs shape public opinion or reinforce dominant narratives.
Approach: They define two types of bias in large language models: model bias and inference bias through a two-phase evaluation.
Outcome: The proposed framework evaluates large language models on factual and disputable questions across four languages and question types.
Bias in the East, Bias in the West: A Bilingual Analysis of LLM Political Bias on U.S.- and China-Related Issues (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) can exhibit political biases, which creates a risk of undue influence on LLM users and public opinion.
Approach: They use a dataset of 36k real-time test prompts to measure LLM political bias on U.S. and Chinese issues.
Outcome: The proposed model origin and prompt language influence bias on 60 political issues.
Inflating Topic Relevance with Ideology: A Case Study of Political Ideology Bias in Social Topic Detection Models (2020.coling-main)

Copied to clipboard

Challenge: a study examines the impact of political ideology biases in training data . topic detection methods may contain or propagate certain biase resulting in a skewed data collection .
Approach: They propose to learn a text representation that is invariant to political ideology while still judging topic relevance.
Outcome: The proposed model can be invariant to political ideology while still judging topic relevance.
OpinionGPT: Modelling Explicit Biases in Instruction-Tuned LLMs (2024.naacl-demo)

Copied to clipboard

Challenge: Current research seeks to de-bias such models, or suppress potentially biased answers.
Approach: They present a web demo to test the biases of instruction-tuned Large Language Models . they identify 11 different biase based on a corpus of data .
Outcome: The proposed demo shows that biases in instruction-tuning are explicit and transparent . the demo shows how the model was trained and showcases the web application .
Assessing Reliability and Political Bias In LLMs’ Judgements of Formal and Material Inferences With Partisan Conclusions (2025.acl-long)

Copied to clipboard

Challenge: This paper examines the ability of LLMs to correctly label simple inferences with partisan conclusions.
Approach: They develop a dataset with formal and material inferences with conclusions that favor either the political left or the political right.
Outcome: The proposed models show that they are unreliable and political bias persists throughout the English and German datasets.
Biased LLMs can Influence Political Decision-Making (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have found that biased LLMs can influence decisions in areas such as medical classifications and educational hiring.
Approach: They conducted two interactive experiments on partisan bias in large language models while completing tasks with either a biased liberal, biased conservative, or unbiased control model.
Outcome: The results show that prior knowledge of AI is weakly correlated with a reduction of the bias, suggesting that AI education can be crucial for mitigating bias effects.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and measures focus on gender and racial biases, but political bias exists in LLMs and can lead to polarization and other harms in downstream applications.
Approach: They propose to analyze the content and style of LLMs generated by political issues and propose a framework that can be scalable to other topics.
Outcome: The proposed framework is easily scalable to other topics and is explainable.
Navigating the Political Compass: Evaluating Multilingual LLMs across Languages and Nationalities (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are ubiquitous in today’s technological landscape, boasting a plethora of applications, and even endangering human jobs in complex and creative fields.
Approach: They evaluate the political bias of 15 multilingual LLMs using the Political Compass Test and assign a nationality to each model.
Outcome: The models on the 50 most populous countries and their official languages exhibit political bias.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations