Probing the Plasticity and Correlation of LLM Value Systems: LLM Value Rankings are Not Stable (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have similar value rankings but little is known about how susceptible they are to external influence and how different values are correlated with each other. |
| Approach: | They propose to use 6 different value transformation prompting methods to examine the plasticity of LLM value systems by comparing them with 8 LLMs. |
| Outcome: | The proposed methods are effective on 8 LLMs and 3 families. |
Similar Papers
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective (2025.findings-acl)
Copied to clipboard
Yipeng Kang, Junqi Wang, Yexin Li, Mengmeng Wang, Wenming Tu, Quansen Wang, Hengli Li, Tingjun Wu, Xue Feng, Fangwei Zhong, Zilong Zheng
| Challenge: | Current approaches to value alignment focus on a few core values, such as helpfulness, harmlessness, and honesty. |
| Approach: | They propose to use latent causal value graphs to guide two lightweight value-steering methods . role-based prompting and sparse autoencoder (SAE) steering are also used . |
| Outcome: | Experiments on Gemma-2B-IT and Llama3-8B- IT show that the proposed methods are effective and controllable. |
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights (2025.acl-long)
Copied to clipboard
| Challenge: | Value-aligned LLMs are more prone to harmful behavior than fine-tuned models . value-aligned models generate text according to the aligned values, which can amplify harmful outcomes. |
| Approach: | They propose to use in-context alignment methods to enhance the safety of value-aligned LLMs. |
| Outcome: | The proposed methods improve value alignment and safety, the authors say . value-aligned models are more prone to harmful behavior than fine-tuned models . |
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios . |
| Approach: | They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs . |
| Outcome: | The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB . |
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing value probing methods that capture in-context information and predict models’ real-world actions are limited and lack systematic comparisons. |
| Approach: | They compare three widely used value probing methods: token likelihood, sequence perplexity, and text generation. |
| Outcome: | The proposed methods exhibit large variances under non-semantic perturbations in prompts and option formats, with sequence perplexity being the most robust overall. |
Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) show remarkable performance across tasks . alignment with human values is critical for their responsible development. |
| Approach: | They propose a framework that evaluates value principles along three desirable properties . they propose supervised fine-tuning, reinforcement learning-based approaches . |
| Outcome: | The proposed framework improves value principles along the three desirable properties of LLMs. |
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models (LLMs) focus on item-level behavioral metrics without capturing how models prioritize competing values as a whole. |
| Approach: | They propose a symmetric human-LLM evaluation framework to measure value-structure alignment . they evaluate 12 LLMs across four model families via 240 replicated Q-sorts . |
| Outcome: | The proposed framework measures value-structure alignment across four model families. |
Internal Value Alignment in Large Language Models through Controlled Value Vector Activation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing LLMs do not possess consistent values, but many have been developed to align them at the behavioral level, including supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). |
| Approach: | They propose a Controlled Value Vector Activation method that directly aligns the internal values of Large Language Models by interpreting how a value is encoded in their latent representations. |
| Outcome: | The proposed method achieves highest success rate across 10 basic values without hurting model performance and fluency, and ensures target values even with opposite and potentially malicious input prompts. |
Inertia in Moral and Value Judgments of Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models behave non-deterministically, and prompting is a common method for steering their outputs. |
| Approach: | They use role-play at scale to study the value orientation and inertia of Large Language Models. |
| Outcome: | The proposed model keeps values skewed in one direction across persona settings. |
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks (2024.acl-long)
Copied to clipboard
| Challenge: | Using benchmarks to evaluate Large Language Models is inconsistent with the assumption that the test prompts within a benchmark represent a random sample from some real-world distribution of interest. |
| Approach: | They propose to use a model's average performance across the test prompts of a benchmark to evaluate its performance. |
| Outcome: | The results show that the correlation between model performance across test prompts and the test prompt can change model rankings on major benchmarks. |
Value Compass Benchmarks: A Comprehensive, Generative and Self-Evolving Platform for LLMs’ Value Evaluation (2025.acl-demo)
Copied to clipboard
Jing Yao, Xiaoyuan Yi, Shitong Duan, Jindong Wang, Yuzhuo Bai, Muhua Huang, Yang Ou, Scarlett Li, Peng Zhang, Tun Lu, Zhicheng Dou, Maosong Sun, James Evans, Xing Xie
| Challenge: | Current evaluation methods for large language models face two key challenges: 1. evaluation validity and 2. Result interpretation reduce the pluralistic and incommensurable values to one-dimensional scores. |
| Approach: | They propose a platform for comprehensive value diagnosis of large language models (LLMs) that provides a generative evaluation paradigm that automatically creates real-world test items co-evolving with ever-advancing LLMs. |
| Outcome: | The proposed platform provides a framework for comprehensive value diagnosis of large language models (LLMs) with fine-grained scores and case studies across 27 value dimensions for 33 leading LLMs, customized comparisons, and visualized analysis of LLM’s alignment with cultural values. |