Challenge: Existing approaches to evaluate cultural alignment of large language models are too trivial and focus on static facts and values.
Approach: They argue for intentionally cultural evaluation: an approach that examines cultural assumptions . they characterize what, how, and circumstances by which culturally contingent considerations arise in evaluation .
Outcome: The authors argue for intentionally cultural evaluation: an approach that examines cultural assumptions embedded in all aspects of evaluation, not just in explicitly cultural tasks.

Similar Papers

Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures.
Approach: They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope.
Outcome: The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches .
Culture is Not Trivia: Sociocultural Theory for Cultural NLP (2025.acl-long)

Copied to clipboard

Challenge: Cultural NLP has experienced rapid growth to meet the need to ensure language technologies are effective and safe across a pluralistic user base.
Approach: They propose to use a well-developed theory of culture to clarify methodological constraints and affordances and offer theoretically-motivated paths forward to achieving cultural competence.
Outcome: The proposed framework clarifies methodological constraints and affordances and offers theoretically-motivated paths forward to achieving cultural competence.
Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory (2026.findings-acl)

Copied to clipboard

Challenge: Recent work in NLP has examined large language models for their understanding of cultural norms across countries, ignoring group consensus or possible multicultural environments.
Approach: They apply cultural consensus theory to the World Values Survey to model multidimensional nuance by ignoring group consensus or over-regularizing consensus.
Outcome: The proposed model misrepresents cultural structures by failing to form cohesive consensus or severely over-regularizing consensus.
Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens (2026.findings-eacl)

Copied to clipboard

Challenge: anthropological accounts of culture often focus on static facts or homogeneous values . large language models are being implemented in translation systems, educational tools and search engines .
Approach: They propose to categorize how benchmarks frame culture such as knowledge, preference, performance, or bias.
Outcome: The proposed framework categorizes how benchmarks frame culture, such as knowledge, preference, performance, or bias.
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: a large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models . however, these methods are constrained and lack nuanced and accurate evaluations based on specific cultural proxies.
Approach: They propose to use the World Values Survey and Hofstede Cultural Dimensions as case studies to examine cultural alignment in Large Language Models.
Outcome: The findings advocate for more robust evaluation frameworks that focus on cultural proxies.
Disentangling Language and Culture for Evaluating Multilingual Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Extensive evaluations of large language models (LLMs) are conducted on a wide range of models, revealing a notable cultural-linguistic synergy phenomenon, where models exhibit better performance when questions are culturally aligned with the language.
Approach: They propose a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of large language models by decomposing evaluation along dimensions of linguistic medium and cultural context.
Outcome: The proposed framework allows for a nuanced analysis of LLMs’ ability to process questions within both native and cross-cultural contexts cross-lingually.
Towards Measuring and Modeling “Culture” in LLMs: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models are biased towards Western, Anglocentric or American cultures, a problem that is arguably detrimental to the performance of LLMs.
Approach: They analyze more than 90 recent papers that aim to study cultural representation and inclusion in large language models.
Outcome: The proposed models are biased towards Western, Anglocentric or American cultures, despite their diversity and their robustness.
FrameNet-Cultures: A Benchmark for Evaluating LLMs via Cross-Cultural Frame Semantics (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation paradigms for large language models lack rigorous methods to evaluate cultural alignment . FRAMENET-CULTURES is an open-ended benchmark for evaluating cultural alignment in LLMs .
Approach: They propose a benchmark for evaluating cultural alignment in large language models based on Fillmore-style frame semantics.
Outcome: The proposed benchmark is based on Fillmore-style frame semantics.
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Reliable multilingual evaluation is difficult and culturally appropriate evaluation is even harder to achieve.
Approach: They propose a multilingual evaluation framework that aims to mitigate these biases by improving translations and annotation practices.
Outcome: The proposed framework improves translation quality and cultural coverage and is culturally sensitive and culturally agnostic.
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness .
Approach: They introduce an automatic multilingual framework for evaluating cultural awareness in large language models across languages, regions, and topics.
Outcome: The framework evaluates open-ended text generation, capturing how models express culturally grounded knowledge in natural language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations