Challenge: Existing benchmarks rely on human annotations that are vulnerable to value-related biases.
Approach: They propose a value portrait benchmark that uses items that capture real-life user-LLM interactions and a rated item based on its similarity to their own thoughts to determine reliability.
Outcome: The proposed framework improves the relevance of assessment results to real-world LLM usage by allowing human subjects to rate items with similarity to their own thoughts and derived correlations between these ratings and the subjects’ actual value scores.

Similar Papers

ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are transforming diverse fields and gaining increasing influence as human proxies.
Approach: They propose a psychometric evaluation pipeline grounded in realistic human-AI interactions to probe value orientations and novel tasks for evaluating value understanding in an open-ended value space.
Outcome: The proposed evaluation pipeline is grounded in realistic human-AI interactions and performs tasks that approximate expert conclusions in value-related extraction and generation tasks.
Value Compass Benchmarks: A Comprehensive, Generative and Self-Evolving Platform for LLMs’ Value Evaluation (2025.acl-demo)

Copied to clipboard

Challenge: Current evaluation methods for large language models face two key challenges: 1. evaluation validity and 2. Result interpretation reduce the pluralistic and incommensurable values to one-dimensional scores.
Approach: They propose a platform for comprehensive value diagnosis of large language models (LLMs) that provides a generative evaluation paradigm that automatically creates real-world test items co-evolving with ever-advancing LLMs.
Outcome: The proposed platform provides a framework for comprehensive value diagnosis of large language models (LLMs) with fine-grained scores and case studies across 27 value dimensions for 33 leading LLMs, customized comparisons, and visualized analysis of LLM’s alignment with cultural values.
Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-Sorts (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models (LLMs) focus on item-level behavioral metrics without capturing how models prioritize competing values as a whole.
Approach: They propose a symmetric human-LLM evaluation framework to measure value-structure alignment . they evaluate 12 LLMs across four model families via 240 replicated Q-sorts .
Outcome: The proposed framework measures value-structure alignment across four model families.
The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models have sparked interest in validating human-like cognitive-behavioral traits.
Approach: They examine whether LLM outputs reflect human-like cognitive-behavioral traits . they find that measuring AOVs embedded within LLMs remains opaque .
Outcome: The proposed model can be used to evaluate human-like cognitive-behavioral traits . the proposed model could be used in writing assistants and other applications .
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have explored personality evaluation of LLMs, but they largely overlook the interplay between culture and personality.
Approach: They propose a large-scale benchmark for evaluating LLMs’ personality expression in culturally grounded, behaviorally rich contexts.
Outcome: The proposed benchmark improves alignment with country-specific human personality distributions and elicits more expressive, culturally coherent outputs compared to existing benchmarks.
Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures.
Approach: They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope.
Outcome: The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches .
Benchmarking Cognitive Biases in Large Language Models as Evaluators (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been shown to be effective as automatic evaluators with simple prompting and in-context learning.
Approach: They assemble 16 Large Language Models and evaluate their outputs by preference ranking . they introduce a cognitive bias benchmark to measure six different cognitive biases in LLM evaluation outputs.
Outcome: The proposed model is biased on the CoBBLer benchmark, indicating that machine preferences are misaligned with humans.
Co-Eval: Augmenting LLM-based Evaluation with Machine Metrics (2025.emnlp-main)

Copied to clipboard

Challenge: Existing LLMs suffer from biases and misalignment due to limited functional understanding and knowledge gaps.
Approach: They introduce a framework that leverages a criteria planner model and optimized machine metrics to enhance the scalability and fairness of LLM-based evaluation.
Outcome: The proposed framework reduces biases and improves alignment with human preferences, with gains of up to 0.324 in Spearman correlation.
METAL: Towards Multilingual Meta-Evaluation (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies show that Large Language Models excel on many standard NLP benchmarks.
Approach: They propose a framework for end-to-end evaluation of Large Language Models as evaluators in multilingual scenarios.
Outcome: The proposed framework evaluates LLMs as evaluators in multilingual scenarios.
Exploring Multilingual Concepts of Human Values in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages? (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior research has revealed that certain abstract concepts are linearly represented as directions in the representation space of LLMs, predominantly centered around English.
Approach: They extend previous research that shows certain abstract concepts are linearly represented as directions in LLMs, predominantly centered around English.
Outcome: The proposed model can be used to align LLMs with human values, and it can generate toxic, untruthful, biased, and even illegal content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations