Challenge: Recent models struggle to show performance over 60%, significantly below the pass mark of the source exams (80%), highlighting the room for improvement.
Approach: They propose to use Korean exams to collect 35,030 questions from an expert-level multiple choice model to capture linguistic and cultural aspects of the Korean language.
Outcome: The proposed benchmark is based on 35,030 questions from original Korean exams.

Similar Papers

CMMLU: Measuring massive multitask language understanding in Chinese (2024.findings-acl)

Copied to clipboard

Challenge: Existing large language models struggle to achieve an accuracy of even 60%, which is the pass mark for Chinese exams.
Approach: They propose to use CMMLU to evaluate Chinese multilingual and Chinese LLMs in a comprehensive benchmark that covers various subjects and settings.
Outcome: The proposed benchmark covers natural sciences, social sciences, engineering, and the humanities and aims to improve on existing models.
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios.
Approach: They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea.
Outcome: The proposed benchmarks represent professional knowledge in Korea.
VMLU Benchmarks: A comprehensive benchmark toolkit for Vietnamese LLMs (2025.acl-long)

Copied to clipboard

Challenge: The evolution of Large Language Models (LLMs) has underscored the need for benchmarks designed for various languages and cultural contexts.
Approach: They propose to use Vietnamese multitask language understanding (VMLU) benchmarks to assess different capabilities of LLMs, including general knowledge, reading comprehension, reasoning, and conversational skills.
Outcome: The VMLU Benchmarks assess LLMs' general knowledge, reading comprehension, reasoning, and conversational skills.
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on linguistic competence of language models with grammatical knowledge.
Approach: They propose to use grammar as a measurable proxy to assess linguistic competence of large language models (LLMs) .
Outcome: The proposed model aims to assess the linguistic competence of large language models (LLMs) and humans in Korean.
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models showcase varied multilingual capabilities across tasks . previous assessments focused on fundamental natural language processing (NLP) or isolated capability-specific tasks.
Approach: They propose a multilingual multitask benchmark to assess multilingual capabilities . they use a large-scale benchmark covering fundamental and capability-specialized datasets .
Outcome: The proposed benchmark compares models and tasks across languages and tasks and examines knowledge transfer from English to other languages.
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages (2025.acl-long)

Copied to clipboard

Challenge: preparing native language MMLU benchmarks is costly and limits representativeness of evaluation datasets.
Approach: They propose to use a Turkic language MMLU benchmark to assess massive multitask language understanding capabilities.
Outcome: The proposed benchmarks are based on a Turkic language morphosyntactic and cultural benchmark . the benchmarks evaluate a diverse range of open and proprietary multilingual large language models .
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing multiple choice question answering benchmarks employ automatic translation for multilingual evaluation, but this approach is error-prone and potentially introduces culturally biased questions.
Approach: They introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU . they evaluate over 20 LLMs including open-source, closed-source and Turkish-adapted models .
Outcome: The proposed benchmarks evaluate the reasoning, comprehension, and mathematical abilities of large language models.
KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Language models are striving to grasp commonsense reasoning, but they are lacking in Korean commons- ense benchmarks.
Approach: They present a fine-grained benchmark dataset focused on Korean commonsense reasoning that includes multiple-choice questions across seven error categories.
Outcome: The proposed datasets show that LLMs struggle with Korean commonsense reasoning . human accuracy benchmarked at approximately 85%, while GPT-4’s performance lags at about 74%, and other LLM models demonstrate an average accuracy of around 42%.
Multi-Dimensional Machine Translation Evaluation: Model Evaluation and Resource for Korean (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on MT evaluation characterize quality of output with a single number . a recent advancement in MT technologies has enabled higher-quality, more nuanced translations .
Approach: They propose a 1200-sentence MQM evaluation benchmark for English-Korean and a reference-free QE setup to evaluate the quality of the translations.
Outcome: The proposed model outperforms the existing model in style and accuracy.
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages other than English.
Approach: They propose a benchmark for assessing the Korean legal language understanding of LLMs consisting of 7 legal knowledge tasks and 4 legal reasoning tasks.
Outcome: The proposed model passes the Uniform Bar Exam in the U.S. but its performance is limited for non-standardized tasks and tasks in languages other than English.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations