Challenge: Existing approaches to evaluate latent values and opinions in large language models suffer from three notable shortcomings.
Approach: They propose to analyze 156k LLM responses to 62 propositions of the Political Compass Test (PCT) generated by 6 LLMs using 420 prompt variations.
Outcome: The proposed analysis of 156k LLM responses to the Political Compass Test (PCT) generated by 6 LLMs shows that tropes are recurrent and consistent across prompts.

Similar Papers

Quantifying the Influence of Irrelevant Contexts on Political Opinions Produced by LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Recent studies have examined the generation of large language models (LLMs) on subjective topics such as political opinions and attitudinal questionnaires.
Approach: They use a Political Compass Test questionnaire to quantify how irrelevant information can systematically bias model opinions in specific directions.
Outcome: The results show that even seemingly unrelated contexts alter model responses in predictable ways.
Evaluating Large Language Model Biases in Persona-Steered Generation (2024.findings-acl)

Copied to clipboard

Challenge: a recent wave of powerful new large language models has raised concerns that their expressed opinions may be biased towards certain political, national or moral viewpoints.
Approach: They define an incongruous persona as a persona with multiple traits where one trait makes its other traits less likely in human survey data.
Outcome: The results show that LLMs are less steerable towards incongruous personas than congruous ones . the models that are fine-tuned with RLHF are more steerable, especially towards stances associated with political liberals and women .
Navigating the Political Compass: Evaluating Multilingual LLMs across Languages and Nationalities (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are ubiquitous in today’s technological landscape, boasting a plethora of applications, and even endangering human jobs in complex and creative fields.
Approach: They evaluate the political bias of 15 multilingual LLMs using the Political Compass Test and assign a nationality to each model.
Outcome: The models on the 50 most populous countries and their official languages exhibit political bias.
Identifying Fine-grained Forms of Populism in Political Discourse: A Case Study on Donald Trump’s Presidential Campaigns (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models excel in a wide range of instruction-following tasks, but their grasp of social science concepts remains underexplored.
Approach: They evaluate pre-trained large language models to identify populist discourse . they use a RoBERTa classifier to analyze campaign speeches by Donald Trump .
Outcome: The proposed model outperforms all new-era instruction-tuned LLMs on populist discourse analysis.
AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts (2025.emnlp-main)

Copied to clipboard

Challenge: Distinguishing LLM-generated text from human-written is a key challenge for safe and ethical NLP, especially in high-stake settings such as persuasive online discourse.
Approach: They propose to use general-purpose linguistic features and domain-specific features related to argument quality to compare human- and LLM-authored arguments.
Outcome: The proposed framework compares arguments by humans and three LLMs using two easily-interpretable feature sets.
Large Language Models Still Exhibit Bias in Long Text (2025.findings-acl)

Copied to clipboard

Challenge: Existing fairness benchmarks for large language models focus on simple tasks . a new framework evaluates biases in LLMs through essay-style prompts .
Approach: They propose a framework that evaluates biases in large language models through essay-style prompts.
Outcome: The proposed framework uncovers subtle biases difficult to detect in simple responses.
Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responses (2025.findings-emnlp)

Copied to clipboard

Challenge: large language models (LLMs) increasingly assist subjective decision-making . prior work uses aggregate human judgments, but demographic variation and its linguistic drivers remain underexplored.
Approach: They analyze how demographic background and empathy level correlate with LLM-generated dilemma responses . they also identify markers that predict group-level differences .
Outcome: The authors show that demographic background and empathy level correlate with LLM preferences . their findings highlight the need for demographically informed LLM evaluations.
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)

Copied to clipboard

Challenge: Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies.
Approach: They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Outcome: The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes.
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks and measures focus on gender and racial biases, but political bias exists in LLMs and can lead to polarization and other harms in downstream applications.
Approach: They propose to analyze the content and style of LLMs generated by political issues and propose a framework that can be scalable to other topics.
Outcome: The proposed framework is easily scalable to other topics and is explainable.
Dissecting Human and LLM Preferences (2024.acl-long)

Copied to clipboard

Challenge: a recent study shows that human and Large Language Model preferences are important for model fine-tuning and evaluation.
Approach: They dissect the preferences of human and 32 different Large Language Models to understand their quantitative composition.
Outcome: The proposed model is compared with 32 different large language models using real-world user-model conversations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations