GIMMICK: Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on Large Vision-Language Models (LVLMs) focus on a narrow range of cultures, focus on only a small number of cultural aspects or evaluate a limited selection of models on ONE task only. |
| Approach: | They propose a multimodal benchmark to assess a broad spectrum of cultural knowledge across 144 countries representing six global macro-regions. |
| Outcome: | The proposed benchmark examines cultural knowledge across 144 countries across six global macro-regions. |
Similar Papers
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)
Copied to clipboard
Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki, Omer Nacar, Youssef Nafea, Safaa Taher Abdelfadil, Abdulfattah Mohammed Yahya, Hamzah Luqman, Nada Almarwani, Samah Aloufi, Baraah Qawasmeh, Houdaifa Atou, Serry Sibaee, Hamzah A. Alsayadi, Walid Al-Dhabyani, Maged S. Al-shaibani, Aya El aatar, Nour Qandos, Rahaf Alhamouri, Samar Ahmad, Mohammed Anwar AL-Ghrawi, Aminetou Yacoub, Ruwa AbuHweidi, Vatimetou Mohamed Lemin, Reem Abdel-Salam, Ahlam Bashiti, Adel Ammar, Aisha Alansari, Ahmed Ashraf, Nora Alturayeif, Alcides Alcoba Inciarte, AbdelRahim A. Elmadany, Mohamedou Cheikh Tourad, Ismail Berrada, Mustafa Jarrar, Shady Shehata, Muhammad Abdul-Mageed
| Challenge: | Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. |
| Approach: | They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. |
| Outcome: | The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X). |
M5 – A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models and their multimodal counterparts have shown significant performance disparities across different languages and cultural contexts. |
| Approach: | They propose to evaluate LLMs on diverse vision-language tasks within a multilingual and multicultural context using M5 benchmark. |
| Outcome: | The proposed benchmarks highlight task-agnostic performance disparities between languages and cultural contexts. |
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model (2025.emnlp-main)
Copied to clipboard
Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin’ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan
| Challenge: | Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content. |
| Approach: | They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set . |
| Outcome: | The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu. |
Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)
Copied to clipboard
Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy, Sjoerd Steenkiste, Lisa Hendricks, Karolina Stanczak, Aishwarya Agrawal
| Challenge: | Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning. |
| Approach: | They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents. |
| Outcome: | The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions. |
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation (2025.naacl-long)
Copied to clipboard
Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa
| Challenge: | Using culture-agnostic subsets, performance drops in many LMMs when evaluated in Japanese. |
| Approach: | They introduce a Japanese benchmark to evaluate large multimodal models on expert-level tasks based on the Japanese cultural context. |
| Outcome: | The proposed benchmark enables comparisons with other benchmarks in other languages based on cultural contexts. |
BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models (2026.eacl-long)
Copied to clipboard
Bryan Chen Zhengyu Tan, Weihua Zheng, Zhengyuan Liu, Nancy F. Chen, Hwaran Lee, Kenny Tsu Wei Choo, Roy Ka-Wei Lee
| Challenge: | Existing evaluations assess static recall or isolated visual grounding, leaving unanswered whether VLMs possess robust and transferable cultural understanding. |
| Approach: | They propose a multimodal, multicultural benchmark to evaluate the robustness of everyday cultural knowledge in vision-language models across linguistic rephrasings and visual modalities. |
| Outcome: | ‘BLEnD-Vis‘ constructs 313 culturally grounded question templates spanning 16 regions and generates three aligned multiple-choice formats. |
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Using large vision-language models to understand cultural contexts is a critical area of research. |
| Approach: | They conduct a thorough evaluation of multimodal models at different scales, focusing on their alignment with cultural values. |
| Outcome: | The proposed models show that they exhibit sensitivity to cultural values but their performance is highly context-dependent. |
Grounding Multilingual Multimodal LLMs With Cultural Knowledge (2025.emnlp-main)
Copied to clipboard
| Challenge: | a new data-centric approach could address cultural gaps in multimodal large language models . despite being trained on billions of image-text pairs, today's models are biased towards English and Western data. |
| Approach: | They propose a data-centric approach that directly grounds MLLMs in cultural knowledge. |
| Outcome: | The proposed approach outperforms open-source models on cultural-focused benchmarks without degrading results on mainstream vision–language tasks. |
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)
Copied to clipboard
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
| Challenge: | Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed . |
| Approach: | They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities. |
| Outcome: | The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech . |
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala (2025.emnlp-main)
Copied to clipboard
Ashmari Pramodya, Nirasha Nelki, Heshan Shalinda, Chamila Liyanage, Yusuke Sakai, Randil Pushpananda, Ruvan Weerasinghe, Hidetaka Kamigaito, Taro Watanabe
| Challenge: | Large Language Models (LLMs) have been evaluated mostly on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. |
| Approach: | They evaluate 26 Large Language Models using a multiple-choice question answering benchmark for Sinhala. |
| Outcome: | The new benchmarks show that Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies, but overall performance remains limited. |