Challenge: Large reasoning models trained to reason explicitly in the verbal space have shown superior performance over general large language models (Guo et al., 2025).
Approach: They propose to use Korean Canonical Legal Benchmark to assess language models' legal reasoning capabilities independently of domain-specific knowledge.
Outcome: The proposed benchmark outperforms general-purpose models in a systematic evaluation of 30+ models.

Similar Papers

Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages other than English.
Approach: They propose a benchmark for assessing the Korean legal language understanding of LLMs consisting of 7 legal knowledge tasks and 4 legal reasoning tasks.
Outcome: The proposed model passes the Uniform Bar Exam in the U.S. but its performance is limited for non-standardized tasks and tasks in languages other than English.
UCL-Bench: A Chinese User-Centric Legal Benchmark for Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing legal benchmarks focusing on knowledge and logic evaluate LLMs on various tasks in legal domain, but few have explored the practical application of LLM by actual users.
Approach: They propose a Chinese user-centric legal benchmark that aims to assess the practical application of LLMs by real users.
Outcome: The proposed model outperforms existing models on various tasks in legal domain but does not outperfect ChatGPT.
PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning.
Approach: They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios.
Outcome: The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics.
From KMMLU-Redux to Pro: A Professional Korean Benchmark Suite for LLM Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Using Korean expert-level benchmarks, Large Language Models can be developed in real-world scenarios.
Approach: They introduce two Korean expert-level benchmarks that reflect professional knowledge in Korea.
Outcome: The proposed benchmarks represent professional knowledge in Korea.
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on linguistic competence of language models with grammatical knowledge.
Approach: They propose to use grammar as a measurable proxy to assess linguistic competence of large language models (LLMs) .
Outcome: The proposed model aims to assess the linguistic competence of large language models (LLMs) and humans in Korean.
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable performances in general domains and are now extending into the expert domain of law.
Approach: They propose a Korean Benchmark for Legal EXplainable QA (KoBLEX) that evaluates provision-grounded, multi-hop legal reasoning.
Outcome: The proposed method outperforms baselines and shows a high correlation with human judgments.
LAiW: A Chinese Legal Large Language Models Benchmark (2025.coling-main)

Copied to clipboard

Challenge: Xie et al., 2023) show that large language models (LLMs) can generate legal text, but lack the legal syllogism . legal experts are cautious about their practical application due to the opaque nature of the LLMs.
Approach: They propose a Chinese legal LLM benchmark structured around the legal syllogism . they evaluate LLMs across three levels of capability, each reflecting a more complex stage of legal .
Outcome: The proposed benchmark identifies that LLMs lack the legal syllogism, which hinders trust and understanding from legal experts.
JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing legal benchmarks evaluate isolated tasks or exam-style questions, failing to capture the procedural interdependencies and adjudicative rigor inherent in professional practice.
Approach: They propose a vertical, depth-oriented, domain-specific benchmark to evaluate Large Language Models (LLMs) in Chinese civil litigation.
Outcome: The proposed benchmarks show that large language models exhibit an "illusion of competence" the results highlight a critical gap between fluent linguistic output and judicial reliability .
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond (2025.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show that Legal-R1 delivers competitive performance across diverse tasks.
Approach: They propose to evaluate 12 large language models across 17 legal tasks across statutory and case-law traditions to determine their general reasoning performance.
Outcome: The proposed model performs well across 17 legal tasks across statutory and case-law traditions.
Pub-LawBench: Public-Oriented Benchmarking for LegalAI (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on legal professionals, not legal professionals.
Approach: They propose a public-oriented LegalAI benchmark grounded in legal functionalism and genre analysis to address this gap.
Outcome: The proposed model evaluates 17 large language models on Pub-LawBench using simple prompts and Chain-of-Thought under a vanilla inference setting.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations