Challenge: Existing evaluation methods fail to probe fragile reasoning capabilities of large language models (LLMs) a single undetected defect can have catastrophic consequences, highlighting the need for a new class of benchmarks to stress-test their reliability against nuanced contractual flaws.
Approach: They propose a benchmark to evaluate the fragility of an LLM’s legal reasoning by producing over 7500 real-world perturbed contracts from foundational datasets like CUAD and ContractNLI.
Outcome: The proposed benchmark evaluates LLMs' ability to detect and reason about fine-grained discrepancies from over 7500 real-world perturbed contracts from foundational datasets like CUAD and ContractNLI.

Similar Papers

The Lawyer That Never Thinks: Consistency and Fairness as Keys to Reliable AI (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used in high-stakes domains like law and research.
Approach: They evaluate six leading Large Language Models on rationality, stability, and ethical fairness through reasoning tests, legal challenges, and bias-sensitive scenarios.
Outcome: The models perform well on reasoning tests, legal challenges, and bias-sensitive scenarios.
PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning.
Approach: They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios.
Outcome: The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics.
CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: a new benchmark is designed to evaluate LLMs on Chinese legal knowledge and its application in reasoning . general pre-training that ingests legal texts without specialized focus compromises reliability of LLM responses . achieving trustworthy legal reasoning in LLM requires a robust synergy of accurate knowledge retrieval and strong general reasoning capabilities.
Approach: They propose a benchmark specifically engineered to evaluate LLMs on Chinese legal knowledge and its application in reasoning.
Outcome: The proposed benchmark evaluates LLMs on Chinese legal knowledge and its application in reasoning.
Evaluating Test-Time Scaling LLMs for Legal Reasoning: OpenAI o1, DeepSeek-R1, and Beyond (2025.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show that Legal-R1 delivers competitive performance across diverse tasks.
Approach: They propose to evaluate 12 large language models across 17 legal tasks across statutory and case-law traditions to determine their general reasoning performance.
Outcome: The proposed model performs well across 17 legal tasks across statutory and case-law traditions.
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) offer new opportunities to enhance the annotation process, particularly for detecting label errors in existing datasets.
Approach: They propose to use an ensemble of large language models to flag mislabeled examples by using an LLM-as-a-judge approach to detect label errors in existing datasets.
Outcome: The proposed method improves label accuracy and consistency in large language models.
JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing legal benchmarks evaluate isolated tasks or exam-style questions, failing to capture the procedural interdependencies and adjudicative rigor inherent in professional practice.
Approach: They propose a vertical, depth-oriented, domain-specific benchmark to evaluate Large Language Models (LLMs) in Chinese civil litigation.
Outcome: The proposed benchmarks show that large language models exhibit an "illusion of competence" the results highlight a critical gap between fluent linguistic output and judicial reliability .
Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to large language models focus on semantic similarity, neglecting the intricate logical structures and reasoning essential for addressing complex legal issues.
Approach: They propose a Logical-Semantic Integration Model (LSIM) that bridges semantic and logical coherence and a supervised framework that integrates semantic features with in-context learning.
Outcome: The proposed framework significantly improves accuracy and reliability on a real-world legal QA dataset.
SummEdits: Measuring LLM Ability at Factual Reasoning Through The Lens of Summarization (2023.emnlp-main)

Copied to clipboard

Challenge: Existing factual consistency benchmarks are inadequate to detect factual inconsistencies in LLMs.
Approach: They propose a protocol for inconsistency detection benchmark creation and implement it in a 10-domain benchmark called SummEdits.
Outcome: The proposed method is 20 times more cost-effective per sample and highly reproducible, as it estimates inter-annotator agreement at about 0.9.
LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for legal general intelligence (GI) are result-oriented and do not evaluate the legal intelligence of large language models (LLMs).
Approach: They propose a Chinese legal benchmark for evaluating legal GI in large language models . they use recent legal cases and exam questions to create multiple-choice questions .
Outcome: The proposed benchmarks lack a systematic evaluation of the legal intelligence of large language models (LLMs) the results show that even the best LLMs lagging behind human legal professionals.
RiddleBench: A New Generative Reasoning Benchmark for LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) show remarkable capabilities, but complex reasoning skills require deeper investigation.
Approach: They propose a benchmark of 1,737 puzzles to test reasoning beyond simple pattern matching.
Outcome: The proposed model performs poorly when faced with reordered constraints or irrelevant information.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations