Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLU (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on large language models based on English datasets do not provide adequate data for evaluating their capabilities beyond English. |
| Approach: | They propose a multi-task language understanding benchmark for Indonesian culture and languages . it measures language proficiency, reasoning abilities and real-world knowledge . |
| Outcome: | The proposed model passes the primary school level in Indonesia, while other models perform at lower levels. |
Similar Papers
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia (2025.naacl-industry)
Copied to clipboard
| Challenge: | Using the entire dataset, shuffling answer options introduces instability in the insurance and finance sectors. |
| Approach: | They propose a dataset for evaluation of performance in vocational and professional certification exams in Indonesia. |
| Outcome: | The proposed dataset includes 8,834 multiple-choice questions from 27 large language models across six key sectors. |
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multiple choice question answering benchmarks employ automatic translation for multilingual evaluation, but this approach is error-prone and potentially introduces culturally biased questions. |
| Approach: | They introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU . they evaluate over 20 LLMs including open-source, closed-source and Turkish-adapted models . |
| Outcome: | The proposed benchmarks evaluate the reasoning, comprehension, and mathematical abilities of large language models. |
IndoCL: Benchmarking Indonesian Language Development Assessment (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent interest has surged in applying natural language processing (NLP) and machine learning (ML) to evaluate language development in both first (L1) and second (L2) language acquisition. |
| Approach: | They propose to use an Indonesian corpus as a benchmark for LDA tasks and to use existing large-scale language models to improve performance. |
| Outcome: | The proposed model extracts language-independent features, relieving laborious computation and reliance on specific language. |
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding (2020.aacl-main)
Copied to clipboard
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, Ayu Purwarianti
| Challenge: | Despite the availability of data on Indonesian, progress on this language is slow . available datasets are scattered, with a lack of documentation and minimal community engagement. |
| Approach: | They propose a resource for training, evaluation, and benchmarking on Indonesian natural language understanding tasks. |
| Outcome: | The proposed resource includes 12 tasks ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity. |
LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages (2025.emnlp-main)
Copied to clipboard
| Challenge: | LORAXBENCH is a benchmark for low-resource languages of Indonesia . it covers reading comprehension, open domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 languages. |
| Approach: | They propose a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks: reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering. |
| Outcome: | The proposed benchmark covers reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 Indonesian languages. |
CMMLU: Measuring massive multitask language understanding in Chinese (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing large language models struggle to achieve an accuracy of even 60%, which is the pass mark for Chinese exams. |
| Approach: | They propose to use CMMLU to evaluate Chinese multilingual and Chinese LLMs in a comprehensive benchmark that covers various subjects and settings. |
| Outcome: | The proposed benchmark covers natural sciences, social sciences, engineering, and the humanities and aims to improve on existing models. |
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP (2020.coling-main)
Copied to clipboard
| Challenge: | despite being spoken by 200 million people, the Indonesian language is underrepresented in NLP research. |
| Approach: | They propose a dataset for Indonesian that includes seven NLP tasks . they also propose 'indonesian language evaluation Montage' tasks that are based on previous work . |
| Outcome: | The proposed dataset shows that IndoBERT outperforms IndoLEM over most of the tasks. |
MalayMMLU: A Multitask Benchmark for the Low-Resource Malay Language (2024.findings-emnlp)
Copied to clipboard
Soon Poh, Sze Jue Yang, Jeraelyn Tan, Lawrence Chieng, Jia Tan, Zhenyu Yu, Foong Mun, Chee Seng Chan
| Challenge: | Large Language Models (LLMs) and Large Vision Language Model (LVLMs) exhibit advanced proficiency in language reasoning and comprehension across a wide array of languages. |
| Approach: | They propose to use a multitask language understanding benchmark specifically designed for the Malay language to assess their proficiency. |
| Outcome: | The proposed model performs well in well-resourced languages, but in low-resource languages such as Bahasa Melayu, they are less studied due to a lack of studies and benchmarks. |
IRUEX: A Study on Large Language Models Problem-Solving Skills in Iran’s University Entrance Exam (2025.coling-main)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have profound implications for education. |
| Approach: | They present a novel multiple-choice educational resource specifically designed to evaluate the performance of Large Language Models (LLMs) they use a dataset that contains 868 questions and 36,485 additional questions . |
| Outcome: | The IRUEX dataset contains 868 questions and 36,485 additional questions. |
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Recognizing LLMs’ capability to generate educational content can lead to advances in automated and personalized learning. |
| Approach: | They propose to evaluate the questioning capability in education as a teacher of large language models by evaluating their generated educational questions. |
| Outcome: | The proposed model can generate educational content that aligns with human perspectives and is more apt as an interdisciplinary teacher. |