| Challenge: | a well-formulated benchmark allows objective and precise evaluation of diverse models. |
| Approach: | They propose a benchmark for Korean balanced evaluation of significant tasks that requires advanced Korean linguistic knowledge. |
| Outcome: | The proposed benchmarks are based on five Korean-language downstream tasks . the data is annotated by humans and thoroughly reviewed to guarantee high data quality. |
Similar Papers
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models (2024.lrec-main)
Copied to clipboard
Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jae cheol Lee, Je Won Yeom, Jihyu Jung, Jung woo Kim, Songseong Kim
| Challenge: | Existing evaluation tools rely on translations of English datasets or translation-specific benchmarks such as WMT 21 to assess large language models. |
| Approach: | They propose a dataset curated to challenge models lacking Korean cultural and contextual depth. |
| Outcome: | The HAE-RAE Bench challenges models lacking Korean cultural and contextual depth by highlighting their aptitude for recalling Korean-specific knowledge and cultural contexts. |
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)
Copied to clipboard
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark (2024.acl-long)
Copied to clipboard
Chanjun Park, Hyeonwoo Kim, Dahyun Kim, SeongHwan Cho, Sanghoon Kim, Sukyung Lee, Yungi Kim, Hwalsuk Lee
| Challenge: | Existing benchmarks for evaluating Large Language Models are limited to the English language. |
| Approach: | They introduce the Open Ko-LLM Leaderboard and Ko-H5 Benchmark as tools for evaluating Large Language Models in Korean using private test sets. |
| Outcome: | The proposed evaluation framework is well integrated in the Korean LLM community. |
Evaluating Large Language Models with Enterprise Benchmarks (2025.naacl-industry)
Copied to clipboard
Bing Zhang, Mikio Takeuchi, Ryo Kawahara, Shubhi Asthana, Maruf Hossain, Guang-Jie Ren, Kate Soule, Yifan Mai, Yada Zhu
| Challenge: | Existing benchmarks lack domain-specific datasets for evaluating large language models . existing benchmarks often lack domain specific datasets, which can be difficult to convert to standardized metrics or regulatory issues. |
| Approach: | They propose to use 25 publicly available domain-specific English benchmarks from diverse domains . they propose to combine a wide range of natural language processing tasks for holistic evaluation . |
| Outcome: | The proposed framework includes 25 publicly available domain-specific English benchmarks from diverse enterprise domains like financial services, legal, climate, cyber security, and 2 public Japanese finance benchmarks. |
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios. |
| Approach: | They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea. |
| Outcome: | The proposed benchmarks represent professional knowledge in Korea. |
F-Eval: Asssessing Fundamental Abilities with Refined Evaluation Methods (2024.acl-long)
Copied to clipboard
Yu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li, Qipeng Guo, Hang Yan, Xipeng Qiu, Xuanjing Huang, Dahua Lin
| Challenge: | Large language models (LLMs) have been evaluated for their instruction-following capabilities but lack references to their fundamental abilities. |
| Approach: | They propose a bilingual evaluation benchmark to evaluate the fundamental abilities of large language models including expression, commonsense and logic. |
| Outcome: | The proposed evaluation methods show higher correlation coefficients and larger distinction than other evaluators. |
FINKRX: Establishing Best Practices for Korean Financial NLP (2025.acl-industry)
Copied to clipboard
| Challenge: | Existing tools to evaluate large language models in the financial domain are limited by the inherently closed nature of the financial industry. |
| Approach: | They present the first open leaderboard for evaluating Korean large language models focused on finance. |
| Outcome: | The proposed model is FINKRX, a fully open and transparent LLM built using these best practices. |
KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Language models are striving to grasp commonsense reasoning, but they are lacking in Korean commons- ense benchmarks. |
| Approach: | They present a fine-grained benchmark dataset focused on Korean commonsense reasoning that includes multiple-choice questions across seven error categories. |
| Outcome: | The proposed datasets show that LLMs struggle with Korean commonsense reasoning . human accuracy benchmarked at approximately 85%, while GPT-4’s performance lags at about 74%, and other LLM models demonstrate an average accuracy of around 42%. |
FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation (2026.findings-eacl)
Copied to clipboard
Yulia Otmakhova, Thinh Hung Truong, Rahmad Mahendra, Zenan Zhai, Rongxin Zhu, Daniel Beck, Jey Han Lau
| Challenge: | FLUKE introduces controlled variations across linguistic levels and leverages large language models with human validation to generate modifications. |
| Approach: | They propose a framework for assessing model robustness through systematic minimal variations of test data. |
| Outcome: | The proposed framework evaluates models and LLMs across six diverse NLP tasks and shows that they are more robust to natural, fluent modifications than base models. |
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) are restricted to high- or mid-resource languages, and evaluate performance on higher-order tasks in reasoning and generation. |
| Approach: | They propose a multilingual benchmarking tool to evaluate lexical comprehension and generation abilities of large language models. |
| Outcome: | The proposed benchmarks cover 2700+ languages and surpasses existing benchmarks in terms of language coverage. |