Challenge: Large Language Models behave non-deterministically, and prompting is a common method for steering their outputs.
Approach: They use role-play at scale to study the value orientation and inertia of Large Language Models.
Outcome: The proposed model keeps values skewed in one direction across persona settings.

Similar Papers

Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective (2025.findings-acl)

Copied to clipboard

Challenge: Current approaches to value alignment focus on a few core values, such as helpfulness, harmlessness, and honesty.
Approach: They propose to use latent causal value graphs to guide two lightweight value-steering methods . role-based prompting and sparse autoencoder (SAE) steering are also used .
Outcome: Experiments on Gemma-2B-IT and Llama3-8B- IT show that the proposed methods are effective and controllable.
Persona Prompting as a Lens on LLM Social Reasoning (2026.eacl-long)

Copied to clipboard

Challenge: Persona prompting (PP) is increasingly used to steer large language models towards user-specific generation, but its effect on rationales remains underexplored.
Approach: They examine how LLM-generated rationales vary when conditioned on different demographic personas . they use word-level rationale annotations to measure agreement with human annotations based on PP .
Outcome: The proposed model improves classification on the most subjective task, but fails to align with real-world demographic counterparts.
Context-Value-Action Architecture for Value-Driven Large Language Model Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLMs exhibit behavioral rigidity, a flaw often masked by the self-referential bias of current "LLM-as-a-judge" evaluations.
Approach: They propose a Context-Value-Action architecture that decouples action generation from cognitive reasoning via a Value Verifier trained on authentic human data to explicitly model dynamic value activation.
Outcome: The proposed architecture significantly outperforms baseline models on 1.1 million real-world interaction traces on CVABench.
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios .
Approach: They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs .
Outcome: The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB .
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths (2026.findings-acl)

Copied to clipboard

Challenge: In short-form surveys and psychometric tests, value-related risks and preferences are often underexplored in practical settings.
Approach: They compare short-form responses to long-form outputs to determine whether value preferences align with those expressed in long-term outputs.
Outcome: The proposed method yields only modest gains in the consistency of value expression.
Are Large Language Models Consistent over Value-laden Questions? (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) appear to bias survey answers toward certain values . however, some argue that LLMs are inconsistent to simulate particular values - a recent study .
Approach: They define value consistency as similarity of answers across paraphrases, related questions and multilingual translations of a question to English, Chinese, German, and Japanese.
Outcome: The proposed model is consistent across paraphrases, use-cases, translations, and within a topic.
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models (2025.acl-industry)

Copied to clipboard

Challenge: Large Language Models exhibit subjective preferences, opinions, and beliefs, which may shape their behavior, influence advice and recommendations, and potentially reinforce certain viewpoints.
Approach: They developed a benchmark to assess LLMs’ subjective inclinations across societal, cultural, ethical, and personal domains.
Outcome: The proposed benchmark assesses LLMs’ subjective inclinations across societal, cultural, ethical, and personal domains.
Conservative Bias in Large Language Models: Measuring Relation Predictions (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit pronounced conservative bias in relation extraction tasks, often defaulting to no_relation label when an appropriate option is unavailable.
Approach: They systematically evaluate the trade-off between conservative bias and hallucination in relation extraction tasks by using SBERT and LLM prompts to quantify this effect.
Outcome: The proposed model defaults to no_relation label twice as often as hallucination, resulting in significant information loss when reasoning is not explicitly included in the output.
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment (2026.acl-long)

Copied to clipboard

Challenge: Current alignment paradigms treat "human values" as a monolithic entity, ignoring the fact that many societies are a mosaic of diverse subgroups with distinct and sometimes conflicting values, preferences, and norms.
Approach: They examine whether Large Language Models can emulate distinct cultural values of subgroups . they use a global value survey to examine the value landscape of a multicultural society .
Outcome: The proposed model improves on unseen, out-of-distribution subgroups by 17.4% . the model widens the disparity between subgroup groups when measured by distance-aware metrics.
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies.
Approach: They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space.
Outcome: The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations