Challenge: Existing studies on large language models based on English datasets do not provide adequate data for evaluating their capabilities beyond English.
Approach: They propose a multi-task language understanding benchmark for Indonesian culture and languages . it measures language proficiency, reasoning abilities and real-world knowledge .
Outcome: The proposed model passes the primary school level in Indonesia, while other models perform at lower levels.

Similar Papers

Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia (2025.naacl-industry)

Copied to clipboard

Challenge: Using the entire dataset, shuffling answer options introduces instability in the insurance and finance sectors.
Approach: They propose a dataset for evaluation of performance in vocational and professional certification exams in Indonesia.
Outcome: The proposed dataset includes 8,834 multiple-choice questions from 27 large language models across six key sectors.
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing multiple choice question answering benchmarks employ automatic translation for multilingual evaluation, but this approach is error-prone and potentially introduces culturally biased questions.
Approach: They introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU . they evaluate over 20 LLMs including open-source, closed-source and Turkish-adapted models .
Outcome: The proposed benchmarks evaluate the reasoning, comprehension, and mathematical abilities of large language models.
IndoCL: Benchmarking Indonesian Language Development Assessment (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent interest has surged in applying natural language processing (NLP) and machine learning (ML) to evaluate language development in both first (L1) and second (L2) language acquisition.
Approach: They propose to use an Indonesian corpus as a benchmark for LDA tasks and to use existing large-scale language models to improve performance.
Outcome: The proposed model extracts language-independent features, relieving laborious computation and reliance on specific language.
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding (2020.aacl-main)

Copied to clipboard

Challenge: Despite the availability of data on Indonesian, progress on this language is slow . available datasets are scattered, with a lack of documentation and minimal community engagement.
Approach: They propose a resource for training, evaluation, and benchmarking on Indonesian natural language understanding tasks.
Outcome: The proposed resource includes 12 tasks ranging from single sentence classification to pair-sentences sequence labeling with different levels of complexity.
LORAXBENCH: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages (2025.emnlp-main)

Copied to clipboard

Challenge: LORAXBENCH is a benchmark for low-resource languages of Indonesia . it covers reading comprehension, open domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 languages.
Approach: They propose a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks: reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering.
Outcome: The proposed benchmark covers reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural question answering across 20 Indonesian languages.
CMMLU: Measuring massive multitask language understanding in Chinese (2024.findings-acl)

Copied to clipboard

Challenge: Existing large language models struggle to achieve an accuracy of even 60%, which is the pass mark for Chinese exams.
Approach: They propose to use CMMLU to evaluate Chinese multilingual and Chinese LLMs in a comprehensive benchmark that covers various subjects and settings.
Outcome: The proposed benchmark covers natural sciences, social sciences, engineering, and the humanities and aims to improve on existing models.
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP (2020.coling-main)

Copied to clipboard

Challenge: despite being spoken by 200 million people, the Indonesian language is underrepresented in NLP research.
Approach: They propose a dataset for Indonesian that includes seven NLP tasks . they also propose 'indonesian language evaluation Montage' tasks that are based on previous work .
Outcome: The proposed dataset shows that IndoBERT outperforms IndoLEM over most of the tasks.
MalayMMLU: A Multitask Benchmark for the Low-Resource Malay Language (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) and Large Vision Language Model (LVLMs) exhibit advanced proficiency in language reasoning and comprehension across a wide array of languages.
Approach: They propose to use a multitask language understanding benchmark specifically designed for the Malay language to assess their proficiency.
Outcome: The proposed model performs well in well-resourced languages, but in low-resource languages such as Bahasa Melayu, they are less studied due to a lack of studies and benchmarks.
IRUEX: A Study on Large Language Models Problem-Solving Skills in Iran’s University Entrance Exam (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have profound implications for education.
Approach: They present a novel multiple-choice educational resource specifically designed to evaluate the performance of Large Language Models (LLMs) they use a dataset that contains 868 questions and 36,485 additional questions .
Outcome: The IRUEX dataset contains 868 questions and 36,485 additional questions.
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Recognizing LLMs’ capability to generate educational content can lead to advances in automated and personalized learning.
Approach: They propose to evaluate the questioning capability in education as a teacher of large language models by evaluating their generated educational questions.
Outcome: The proposed model can generate educational content that aligns with human perspectives and is more apt as an interdisciplinary teacher.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations