Challenge: Existing studies on the Open Ko-LLM Leaderboard have been limited to five months . this limited analysis of the Open LLM Leaderboard provides a more comprehensive understanding of the progress in developing large language models.
Approach: They conduct a longitudinal study over eleven months to address limitations of previous studies . they analyze 1,769 models over this period to provide a more comprehensive understanding .
Outcome: The study extends observation period of the Open Ko-LLM Leaderboard to eleven months . primary questions are: What are the specific challenges in improving LLM performance?

Similar Papers

Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs (2025.naacl-industry)

Copied to clipboard

Challenge: Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models . however, the leaderboard has faced significant limitations over time due to its academic nature .
Approach: They propose an improved version of the Open Ko-LLM Leaderboard to improve benchmarking . original benchmarks replaced with new tasks that align with real-world capabilities . four new native Korean benchmarks are introduced to better reflect distinct characteristics of Korean language .
Outcome: The proposed framework improves the Open Ko-LLM Leaderboard2 benchmark suite.
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating Large Language Models are limited to the English language.
Approach: They introduce the Open Ko-LLM Leaderboard and Ko-H5 Benchmark as tools for evaluating Large Language Models in Korean using private test sets.
Outcome: The proposed evaluation framework is well integrated in the Korean LLM community.
From Parameters to Performance: A Data-Driven Study on LLM Structure and Development (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models have revolutionized a wide range of domains, driving significant advancements in both technology and real-world applications.
Approach: They present a large-scale dataset encompassing diverse open-source LLM structures and their performance across multiple benchmarks.
Outcome: The proposed model validates the relationship between structural configurations and performance across multiple benchmarks and further corroborates the findings using mechanistic interpretability techniques.
LongLeader: A Comprehensive Leaderboard for Large Language Models in Long-context Scenarios (2025.naacl-long)

Copied to clipboard

Challenge: LongLeader aims to assess different LLMs' long-context comprehension abilities . long-constext comprehension is a key bottleneck for many use cases .
Approach: They propose a leaderboard to assess different LLMs' long-context comprehension abilities . they offer open-source access to the benchmarks and maintain a dedicated website .
Outcome: The proposed model assesses different LLMs on selected benchmarks and provides open-source access to the benchmarks.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains.
Approach: They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks .
Outcome: The proposed evaluations are reproducible, reliable, and robust.
Assessing the Capabilities of Large Language Models in Coreference: An Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a new approach to coreference resolution, but their performance is not yet fully understood.
Approach: They propose that future efforts should improve scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs.
Outcome: The proposed methods improve scope, data, and evaluation methods of traditional coreference research to adapt to the development of LLMs.
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)

Copied to clipboard

Challenge: Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results.
Approach: They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories.
Outcome: The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories.
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation.
Approach: They propose a generic workflow for LLM-driven synthetic data generation.
Outcome: The proposed workflows highlight gaps in existing research and outline avenues for future studies.
FINKRX: Establishing Best Practices for Korean Financial NLP (2025.acl-industry)

Copied to clipboard

Challenge: Existing tools to evaluate large language models in the financial domain are limited by the inherently closed nature of the financial industry.
Approach: They present the first open leaderboard for evaluating Korean large language models focused on finance.
Outcome: The proposed model is FINKRX, a fully open and transparent LLM built using these best practices.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations