Challenge: Existing studies have evaluated cultural knowledge of large language models, but they fail to assess dynamic cultural competence.
Approach: They propose a benchmark to assess cultural competence through intercultural scenarios that span 60 countries across six continents.
Outcome: The proposed benchmark measures the ability of large language models to apply cultural knowledge effectively in cross-cultural interactions.

Similar Papers

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing multilingual evaluation benchmarks neglect cultural nuances and lack language coverage in subjective tasks.
Approach: They propose a framework that categorizes evaluation tasks into three cultural layers and nine cognitive sub-layers.
Outcome: The proposed framework surpasses prior coverage by up to 111% on 20+ LLMs.
LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly deployed as autonomous agents . evaluations focus primarily on task success rather than cultural appropriateness or reliability.
Approach: They propose a multi-cultural, dynamic benchmark that embeds large language models as agents in a simulated town and evaluates them on task completion and adherence to socio-cultural norms.
Outcome: The proposed model evaluates LLMs on task completion and adherence to socio-cultural norms across models and cultural profiles.
Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies show that Large Language Models are biased towards a Western and Anglo-centric worldview.
Approach: They propose to extend the Octopus test to measure "cultural awareness" they argue that cultural awareness is needed for AI systems to be useful across cultures .
Outcome: The proposed method argues that cultural awareness is not cultural knowledge, but meta-cultural competence . the proposed method is based on the octopus test, which shows it is impossible to learn meaning from real-world concepts without knowing intent and meaning .
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are widely used and engage millions of users from diverse contexts and cultures.
Approach: They propose an evaluation framework to assess LLMs’ cultural adaptability by measuring their ability to judge social acceptability across varying levels of cultural norm specificity.
Outcome: The proposed model shows stronger adaptability to English-centric cultures over those from the Global South.
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios (2025.naacl-long)

Copied to clipboard

Challenge: Large language models are increasingly employed to empower autonomous agents to simulate human behavior.
Approach: They propose to evaluate LLM-driven agents through multi-turn interactions using a bottom-up approach to create diverse social scenarios constructed from extensive scripts.
Outcome: The proposed model evaluates LLM-driven agents through multi-turn interactions emphasizing goal completion and implicit reasoning.
Evaluating Cultural and Social Awareness of LLM Web Agents (2025.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks often overlook cultural and social awareness . current evaluations focus on task completion, often ignoring the diverse cultural and socio-cultural backgrounds.
Approach: They propose a benchmark to assess LLM agents’ sensitivity to cultural and social norms across two web-based tasks: online shopping and social discussion forums.
Outcome: The proposed framework evaluates LLM agents’ ability to detect and appropriately respond to norm-violating user queries and observations across two web-based tasks.
Cultural Learning-Based Culture Adaptation of Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches for adapting large language models to diverse cultural values often rely on prompt engineering.
Approach: They propose a framework for enhancing LLM alignment with cultural values based on cultural learning that leverages simulated social interactions to generate role-playing scenarios.
Outcome: The proposed framework improves cultural value alignment across various model architectures measured using World Value Survey data.
Disentangling Language and Culture for Evaluating Multilingual Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Extensive evaluations of large language models (LLMs) are conducted on a wide range of models, revealing a notable cultural-linguistic synergy phenomenon, where models exhibit better performance when questions are culturally aligned with the language.
Approach: They propose a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of large language models by decomposing evaluation along dimensions of linguistic medium and cultural context.
Outcome: The proposed framework allows for a nuanced analysis of LLMs’ ability to process questions within both native and cross-cultural contexts cross-lingually.
Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures.
Approach: They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope.
Outcome: The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches .
MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness .
Approach: They introduce an automatic multilingual framework for evaluating cultural awareness in large language models across languages, regions, and topics.
Outcome: The framework evaluates open-ended text generation, capturing how models express culturally grounded knowledge in natural language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations