Challenge: a well-formulated benchmark allows objective and precise evaluation of diverse models.
Approach: They propose a benchmark for Korean balanced evaluation of significant tasks that requires advanced Korean linguistic knowledge.
Outcome: The proposed benchmarks are based on five Korean-language downstream tasks . the data is annotated by humans and thoroughly reviewed to guarantee high data quality.

Similar Papers

HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing evaluation tools rely on translations of English datasets or translation-specific benchmarks such as WMT 21 to assess large language models.
Approach: They propose a dataset curated to challenge models lacking Korean cultural and contextual depth.
Outcome: The HAE-RAE Bench challenges models lacking Korean cultural and contextual depth by highlighting their aptitude for recalling Korean-specific knowledge and cultural contexts.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating Large Language Models are limited to the English language.
Approach: They introduce the Open Ko-LLM Leaderboard and Ko-H5 Benchmark as tools for evaluating Large Language Models in Korean using private test sets.
Outcome: The proposed evaluation framework is well integrated in the Korean LLM community.
Evaluating Large Language Models with Enterprise Benchmarks (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks lack domain-specific datasets for evaluating large language models . existing benchmarks often lack domain specific datasets, which can be difficult to convert to standardized metrics or regulatory issues.
Approach: They propose to use 25 publicly available domain-specific English benchmarks from diverse domains . they propose to combine a wide range of natural language processing tasks for holistic evaluation .
Outcome: The proposed framework includes 25 publicly available domain-specific English benchmarks from diverse enterprise domains like financial services, legal, climate, cyber security, and 2 public Japanese finance benchmarks.
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios.
Approach: They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea.
Outcome: The proposed benchmarks represent professional knowledge in Korea.
F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been evaluated for their instruction-following capabilities but lack references to their fundamental abilities.
Approach: They propose a bilingual evaluation benchmark to evaluate the fundamental abilities of large language models including expression, commonsense and logic.
Outcome: The proposed evaluation methods show higher correlation coefficients and larger distinction than other evaluators.
FINKRX: Establishing Best Practices for Korean Financial NLP (2025.acl-industry)

Copied to clipboard

Challenge: Existing tools to evaluate large language models in the financial domain are limited by the inherently closed nature of the financial industry.
Approach: They present the first open leaderboard for evaluating Korean large language models focused on finance.
Outcome: The proposed model is FINKRX, a fully open and transparent LLM built using these best practices.
KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Language models are striving to grasp commonsense reasoning, but they are lacking in Korean commons- ense benchmarks.
Approach: They present a fine-grained benchmark dataset focused on Korean commonsense reasoning that includes multiple-choice questions across seven error categories.
Outcome: The proposed datasets show that LLMs struggle with Korean commonsense reasoning . human accuracy benchmarked at approximately 85%, while GPT-4’s performance lags at about 74%, and other LLM models demonstrate an average accuracy of around 42%.
FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation (2026.findings-eacl)

Copied to clipboard

Challenge: FLUKE introduces controlled variations across linguistic levels and leverages large language models with human validation to generate modifications.
Approach: They propose a framework for assessing model robustness through systematic minimal variations of test data.
Outcome: The proposed framework evaluates models and LLMs across six diverse NLP tasks and shows that they are more robust to natural, fluent modifications than base models.
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are restricted to high- or mid-resource languages, and evaluate performance on higher-order tasks in reasoning and generation.
Approach: They propose a multilingual benchmarking tool to evaluate lexical comprehension and generation abilities of large language models.
Outcome: The proposed benchmarks cover 2700+ languages and surpasses existing benchmarks in terms of language coverage.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations