MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment (2026.findings-eacl)
Copied to clipboard
Omid Ghahroodi, Arshia Hemmat, Marzia Nouri, Seyed Mohammad Hadi Hosseini, Doratossadat Dastgheib, Mohammad Vali Sanian, Alireza Sahebi, Reihaneh Zohrabi, Mohammad Hossein Rohban, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah
| Challenge: | Recent advances in large vision-language models have primarily focused on English, with limited attention given to other languages. |
| Approach: | They propose a dataset to evaluate Persian VLMs across scientific, reasoning, and human-level understanding tasks. |
| Outcome: | The proposed model performs well across scientific reasoning, reasoning, and human-level understanding tasks in Persian and English. |
Similar Papers
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation (2025.findings-naacl)
Copied to clipboard
João Matos, Shan Chen, Siena Kathleen V. Placino, Yingya Li, Juan Carlos Climent Pardo, Daphna Idan, Takeshi Tohyama, David Restrepo, Luis Filipe Nakayama, José María Millet Pascual-Leone, Guergana K Savova, Hugo Aerts, Leo Anthony Celi, An-Kwok Ian Wong, Danielle Bitterman, Jack Gallifant
| Challenge: | Existing multiple-choice question and answer (QA) datasets are text-only and available in a limited subset of languages and countries. |
| Approach: | They propose a multilingual, multimodal benchmarking dataset to evaluate multimodal/vision language models in healthcare. |
| Outcome: | The WorldMedQA-V includes 568 labeled multiple-choice QAs paired with 568 medical images from four countries. |
Matina: A Culturally-Aligned Persian Language Model Using Multiple LoRA Experts (2025.findings-acl)
Copied to clipboard
Sara Bourbour Hosseinbeigi, MohammadAli SeifKashani, Javad Seraj, Fatemeh Taherinezhad, Ali Nafisi, Fatemeh Nadi, Iman Barati, Hosein Hasani, Mostafa Amiri, Mostafa Masoudi
| Challenge: | Existing Large language models fail to accurately model underrepresented languages and cultures, limiting their applicability and acceptance. |
| Approach: | They develop a Persian-focused multi-expert model that incorporates Iranian cultural values and linguistic structures. |
| Outcome: | The proposed model outperforms baseline models in task performance and user satisfaction. |
EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for vision language models are outdated and unable to accurately assess their performance. |
| Approach: | They propose a multi-discipline multimodal multilingual exam benchmark for vision language models . they collect multiple-choice questions across 20 disciplines across 11 languages from 7 language families . |
| Outcome: | The EXAMS-V exam includes 20,932 multiple-choice questions across 20 disciplines . the questions come in 11 languages from 7 language families and require advanced reasoning skills . |
JEEM: Vision-Language Understanding in Four Arabic Dialects (2026.findings-eacl)
Copied to clipboard
Karima Kadaoui, Hanin Atwany, Hamdan Al-Ali, Abdelrahman Mohamed, Ali Mekky, Sergei Tilga, Natalia Fedorova, Ekaterina Artemova, Hanan Aldarmaki, Yova Kementchedjhieva
| Challenge: | Existing evaluation datasets feature Western-centric images and English text, while their non-English counterparts are often derived from the latter. |
| Approach: | They propose to evaluate Vision-Language Models (VLMs) on visual understanding across four Arabic-speaking countries: Jordan, The Emirates, Egypt, and Morocco. |
| Outcome: | The proposed model underperforms in visual understanding and dialect-specific generation across four Arabic-speaking countries. |
EXAMS: A Multi-subject High School Examinations Dataset for Cross-lingual and Multilingual Question Answering (2020.emnlp-main)
Copied to clipboard
| Challenge: | EXAMS is a benchmark dataset for cross-lingual and multilingual question answering for high school examinations. |
| Approach: | They propose to use EXAMS to evaluate cross-lingual and multilingual question answering for high school examinations. |
| Outcome: | The proposed model can be used to explore multilingual reasoning and knowledge transfer methods and pre-trained models in schools in different languages, which was not possible by now. |
IRUEX: A Study on Large Language Models Problem-Solving Skills in Iran’s University Entrance Exam (2025.coling-main)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have profound implications for education. |
| Approach: | They present a novel multiple-choice educational resource specifically designed to evaluate the performance of Large Language Models (LLMs) they use a dataset that contains 868 questions and 36,485 additional questions . |
| Outcome: | The IRUEX dataset contains 868 questions and 36,485 additional questions. |
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT (2024.lrec-main)
Copied to clipboard
Amirhossein Abaskohi, Sara Baruni, Mostafa Masoudi, Nesa Abbasi, Mohammad Hadi Babalou, Ali Edalat, Sepehr Kamahi, Samin Mahdizadeh Sani, Nikoo Naghavian, Danial Namazifard, Pouya Sadeghi, Yadollah Yaghoobzadeh
| Challenge: | a new study examines the efficacy of large language models (LLMs) for Persian . ChatGPT and LLMs have shown remarkable performance in English, but their efficiency for low-resource languages remains an open question. |
| Approach: | They present a benchmarking study of large language models (LLMs) for Persian . they focus on GPT-3.5-turbo, but also GPT-4 and OpenChat-3.5 . |
| Outcome: | The proposed model performs better in Persian than other low-resource languages . the study is the first comprehensive benchmarking of large language models . |
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation (2025.coling-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are mainly trained on English data and struggle with low-resource languages. |
| Approach: | They propose to add a new language to Llama to improve classification accuracy for Persian tasks by aligning representations through bilingual pretraining and instruction datasets. |
| Outcome: | The proposed model performs on generation and classification tasks with no adverse impact and sometimes even improvements on English tasks. |
Matina: A Large-Scale 73B Token Persian Text Corpus (2025.naacl-long)
Copied to clipboard
Sara Bourbour Hosseinbeigi, Fatemeh Taherinezhad, Heshaam Faili, Hamed Baghbani, Fatemeh Nadi, Mostafa Amiri
| Challenge: | Existing Persian datasets are small and lack content diversity . lack of high-quality data has slowed development of NLP models and open-source LLMs for Persian. |
| Approach: | They propose a Persian dataset of 72.9B tokens that is preprocessed and deduplicated to ensure high data quality. |
| Outcome: | The proposed model performs well on key Persian NLP tasks. |
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios (2026.acl-long)
Copied to clipboard
Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin, Siming Liu, Xinyue Liang, Yaolin Li, Zhuangzhi Dong, Jingren Zhang, Yufan Deng, Xinyu Zou, Yang Gao, Heyan Huang
| Challenge: | Existing benchmarks that focus on knowledge-intensive tasks do not reflect diverse educational scenarios. |
| Approach: | They propose a benchmark that incorporates 9 major scenarios and 4,000 educational contexts. |
| Outcome: | The proposed model performs comparable to state-of-the-art large models on the test set. |