CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs’ Cultural Knowledge Through Human-AI Red-Teaming (2025.acl-long)
Copied to clipboard
Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi
| Challenge: | CulturalBench is a set of 1,696 human-written and human-verified questions to assess LMs’ cultural knowledge covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru. |
| Approach: | They construct a set of 1,696 human-written and human-verified questions to assess LMs' cultural knowledge, covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru. |
| Outcome: | The proposed model outperforms other models across cultures, while underperforming on questions related to North Africa, South America and Middle East. |
Similar Papers
WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | a global dataset for multi-cultural value prediction task is lacking in the computer science community . a multi-culture awareness of LMs is critical to generating safe and personalized responses . |
| Approach: | They present a global multi-cultural value prediction task using a world value survey dataset . they construct more than 20 million examples of the type "(demographic attributes, value question) answer" they show that the task is challenging for strong open and closed-source models . |
| Outcome: | The proposed model can generate a rating response to a value question based on demographic contexts on 11.1%, 25.0%, 72.2%, and 75.0% of the questions. |
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Cultural competence is defined as the ability to understand and adapt to multicultural contexts. |
| Approach: | They propose a framework that uses a hierarchical multilingual taxonomy and a Retrieval-Augmented Generation to synthesize culturally relevant question-answer pairs. |
| Outcome: | The proposed framework contains a hierarchical multilingual taxonomy covering 12 primary and 130 secondary topics and a Retrieval-Augmented Generation (RAG)-based methodology leveraging factual knowledge to synthesize culturally relevant question-answer pairs. |
Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) serve diverse global audiences, making it critical for responsible AI deployment across cultures. |
| Approach: | They propose a framework that conceptualizes alignment along three dimensions: Cultural Group, Cultural Elements and Awareness Scope. |
| Outcome: | The proposed framework reveals critical gaps between benchmarks and real-world cultural biases . region dominates cultural group representation, social and political relations dominates coverage . majority of datasets adopt majority-focused Awareness Scope approaches . |
The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models (2026.acl-long)
Copied to clipboard
Yilun Liu, Chunguang Zhao, Mengyao Piao, Lingqi Miao, Shimin Tao, Minggui HE, Chenxin Liu, Zhang Li, null Mahongxia, Jiaxin Guo, Chen Liu, Liqun Deng, Jiansheng Wei, Xiaojun Meng, Fanyi Du, Daimeng Wei, Yanghua Xiao
| Challenge: | Existing multilingual evaluation benchmarks neglect cultural nuances and lack language coverage in subjective tasks. |
| Approach: | They propose a framework that categorizes evaluation tasks into three cultural layers and nine cognitive sub-layers. |
| Outcome: | The proposed framework surpasses prior coverage by up to 111% on 20+ LLMs. |
GlobalBench: A Benchmark for Global Progress in Natural Language Processing (2023.emnlp-main)
Copied to clipboard
Yueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig
| Challenge: | despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages. |
| Approach: | They propose to use global benchmarks to track progress on all NLP datasets in all languages. |
| Outcome: | a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench. |
LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed as autonomous agents . evaluations focus primarily on task success rather than cultural appropriateness or reliability. |
| Approach: | They propose a multi-cultural, dynamic benchmark that embeds large language models as agents in a simulated town and evaluates them on task completion and adherence to socio-cultural norms. |
| Outcome: | The proposed model evaluates LLMs on task completion and adherence to socio-cultural norms across models and cultural profiles. |
VideoVista-CulturalLingo: 360° Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension (2025.acl-long)
Copied to clipboard
| Challenge: | Existing video evaluation benchmarks focus on a single language, typically English, and feature videos rooted in Western cultural contexts. |
| Approach: | They propose a video evaluation benchmark designed to bridge cultural, linguistic, and domain divide in video comprehension. |
| Outcome: | The proposed video evaluation benchmark bridges cultural, linguistic, and domain divides . existing benchmarks only feature videos from YouTube, Shutterstock, or established video datasets based on cultural diversity . |
LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models (2026.acl-long)
Copied to clipboard
Jian Gao, Richeng Xuan, Zhaolu Kang, Dingshi Liao, Wenxin Huang, Zongmou Huang, Yangdi Xu, Bowen Qin, Zheqi He, Xi Yang, null Changjinli, Yonghua Lin
| Challenge: | Existing SEA-focused benchmarks miss Lao-specific cultural grounding and linguistic properties. |
| Approach: | They propose a multi-dimensional benchmark for assessing large language models in Lao . they use open-source and held-out subsets to evaluate languages with a hybrid pipeline . |
| Outcome: | LaoBench is the first large-scale, high-quality, and multidimensional benchmark for assessing LLM language understanding and reasoning in Lao. |
SEA-SafeguardBench: Culturally Grounded Safety Benchmark for Southeast Asian Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing multilingual safety benchmarks rely on machine-translated English data, which fails to capture nuances in low-resource languages. |
| Approach: | They propose to use a human-verified safety benchmark for Southeast Asian languages to validate their safety and cultural diversity. |
| Outcome: | The proposed model outperforms existing models in general, in-the-wild, and content generation across eight languages and 21,640 samples across three subsets: general, and in- the-wild. |
Let’s Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models’ Understanding of Sports (2025.emnlp-main)
Copied to clipboard
Punit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, Jose G Moreno
| Challenge: | Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. |
| Approach: | They propose to use multiple-choice questions (MCQs) to assess LMs' understanding of traditional sports across 60 countries and 6 continents. |
| Outcome: | The new benchmark will be publicly available, fostering research in culturally aware AI systems. |