Challenge: CulturalBench is a set of 1,696 human-written and human-verified questions to assess LMs’ cultural knowledge covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru.
Approach: They construct a set of 1,696 human-written and human-verified questions to assess LMs' cultural knowledge, covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru.
Outcome: The proposed model outperforms other models across cultures, while underperforming on questions related to North Africa, South America and Middle East.

Similar Papers

WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models (2024.lrec-main)

Copied to clipboard

Challenge: a global dataset for multi-cultural value prediction task is lacking in the computer science community . a multi-culture awareness of LMs is critical to generating safe and personalized responses .
Approach: They present a global multi-cultural value prediction task using a world value survey dataset . they construct more than 20 million examples of the type "(demographic attributes, value question) answer" they show that the task is challenging for strong open and closed-source models .
Outcome: The proposed model can generate a rating response to a value question based on demographic contexts on 11.1%, 25.0%, 72.2%, and 75.0% of the questions.
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis (2025.findings-emnlp)

Copied to clipboard

Challenge: Cultural competence is defined as the ability to understand and adapt to multicultural contexts.
Approach: They propose a framework that uses a hierarchical multilingual taxonomy and a Retrieval-Augmented Generation to synthesize culturally relevant question-answer pairs.
Outcome: The proposed framework contains a hierarchical multilingual taxonomy covering 12 primary and 130 secondary topics and a Retrieval-Augmented Generation (RAG)-based methodology leveraging factual knowledge to synthesize culturally relevant question-answer pairs.
Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures.
Approach: They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope.
Outcome: The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches .
The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing multilingual evaluation benchmarks neglect cultural nuances and lack language coverage in subjective tasks.
Approach: They propose a framework that categorizes evaluation tasks into three cultural layers and nine cognitive sub-layers.
Outcome: The proposed framework surpasses prior coverage by up to 111% on 20+ LLMs.
GlobalBench: A Benchmark for Global Progress in Natural Language Processing (2023.emnlp-main)

Copied to clipboard

Challenge: despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages.
Approach: They propose to use global benchmarks to track progress on all NLP datasets in all languages.
Outcome: a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench.
LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly deployed as autonomous agents . evaluations focus primarily on task success rather than cultural appropriateness or reliability.
Approach: They propose a multi-cultural, dynamic benchmark that embeds large language models as agents in a simulated town and evaluates them on task completion and adherence to socio-cultural norms.
Outcome: The proposed model evaluates LLMs on task completion and adherence to socio-cultural norms across models and cultural profiles.
VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension (2025.acl-long)

Copied to clipboard

Challenge: Existing video evaluation benchmarks focus on a single language, typically English, and feature videos rooted in Western cultural contexts.
Approach: They propose a video evaluation benchmark designed to bridge cultural, linguistic, and domain divide in video comprehension.
Outcome: The proposed video evaluation benchmark bridges cultural, linguistic, and domain divides . existing benchmarks only feature videos from YouTube, Shutterstock, or established video datasets based on cultural diversity .
LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing SEA-focused benchmarks miss Lao-specific cultural grounding and linguistic properties.
Approach: They propose a multi-dimensional benchmark for assessing large language models in Lao . they use open-source and held-out subsets to evaluate languages with a hybrid pipeline .
Outcome: LaoBench is the first large-scale, high-quality, and multidimensional benchmark for assessing LLM language understanding and reasoning in Lao.
SEA-SafeguardBench: Culturally Grounded Safety Benchmark for Southeast Asian Languages (2026.findings-acl)

Copied to clipboard

Challenge: Existing multilingual safety benchmarks rely on machine-translated English data, which fails to capture nuances in low-resource languages.
Approach: They propose to use a human-verified safety benchmark for Southeast Asian languages to validate their safety and cultural diversity.
Outcome: The proposed model outperforms existing models in general, in-the-wild, and content generation across eight languages and 21,640 samples across three subsets: general, and in- the-wild.
Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sports (2025.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions.
Approach: They propose to use multiple-choice questions (MCQs) to assess LMs' understanding of traditional sports across 60 countries and 6 continents.
Outcome: The new benchmark will be publicly available, fostering research in culturally aware AI systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations