Challenge: Existing tools to evaluate large language models in the financial domain are limited by the inherently closed nature of the financial industry.
Approach: They present the first open leaderboard for evaluating Korean large language models focused on finance.
Outcome: The proposed model is FINKRX, a fully open and transparent LLM built using these best practices.

Similar Papers

Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs (2025.naacl-industry)

Copied to clipboard

Challenge: Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models . however, the leaderboard has faced significant limitations over time due to its academic nature .
Approach: They propose an improved version of the Open Ko-LLM Leaderboard to improve benchmarking . original benchmarks replaced with new tasks that align with real-world capabilities . four new native Korean benchmarks are introduced to better reflect distinct characteristics of Korean language .
Outcome: The proposed framework improves the Open Ko-LLM Leaderboard2 benchmark suite.
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating Large Language Models are limited to the English language.
Approach: They introduce the Open Ko-LLM Leaderboard and Ko-H5 Benchmark as tools for evaluating Large Language Models in Korean using private test sets.
Outcome: The proposed evaluation framework is well integrated in the Korean LLM community.
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages other than English.
Approach: They propose a benchmark for assessing the Korean legal language understanding of LLMs consisting of 7 legal knowledge tasks and 4 legal reasoning tasks.
Outcome: The proposed model passes the Uniform Bar Exam in the U.S. but its performance is limited for non-standardized tasks and tasks in languages other than English.
Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing financial benchmarks suffer from limited language and task coverage, low-quality datasets, and inadequate adaptability for LLM evaluation.
Approach: They propose a bilingual benchmark for financial LLMs that assesses models’ language understanding and generation capabilities.
Outcome: The proposed bilingual benchmark assesses models’ language understanding and generation capabilities.
FinMaster: A Holistic Benchmark for Full-Pipeline Financial Management with Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks lack domain-specific data, realistic workflow-level task design, and standardized workflow- level evaluation.
Approach: a new benchmark evaluates large language models on financial management workflows . the global financial services market is projected to grow to $37 trillion by 2027 .
Outcome: a new benchmark for large language models on financial management workflows reveals critical capability gaps . accuracy drops from 90% on basic tasks to 40% on complex scenarios requiring multi-step reasoning . the global financial services market reached $25.8 trillion in 2022 and is projected to grow to $37 trillion by 2027 .
CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have achieved remarkable performance on various NLP tasks, yet their potential in more challenging task like finance, has not been fully explored.
Approach: They propose a benchmark to assess the financial knowledge of large language models (LLMs) in China.
Outcome: The proposed benchmark is the most comprehensive evaluation benchmark to date for LLMs in finance.
Large Language Model Agents in Finance: A Survey Bridging Research, Practice, and Real-World Deployment (2025.findings-emnlp)

Copied to clipboard

Challenge: a systematic review of large language models (LLMs) is conducted to better align their capabilities with real-world demands.
Approach: They propose a functional taxonomy mapping financial domains to tasks, datasets, and institutional constraints. they catalog over 30 financial benchmarks and 20 representative models.
Outcome: The proposed model frameworks are bridging financial practice and LLM research.
When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive performance on a variety of tasks and domains.
Approach: They propose a domain specific financial LANGuage model which uses financial keywords and phrases for better masking.
Outcome: The proposed model outperforms existing models on a variety of tasks and domains.
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks, but their proficiency and reliability in the specialized domain of financial data analysis remain uncertain.
Approach: FinDABench is a benchmark designed to evaluate the financial data analysis capabilities of Large Language Models (LLMs) it comprises 15,200 training instances and 8,900 test instances, all meticulously crafted by human experts.
Outcome: FinDABench measures the financial data analysis capabilities of large language models (LLMs) across three dimensions: 1) Core Ability; 2) Analytical Ability; 3) Technical Ability.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations