Challenge: Contract review is labor-intensive, time-consuming, and costly . a benchmark is proposed to detect potential legal conflicts .
Approach: They propose a benchmark for legal provision recommendation and conflict detection for contract auto-reviewing which aims to recommend the legal provisions related to contract clauses and detect possible legal conflicts.
Outcome: The proposed task recommends legal provisions related to contract clauses and detects legal conflicts.

Similar Papers

Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts (2024.lrec-main)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) as foundational models, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts.
Approach: They propose to use a new annotation scheme to classify clauses in Terms-and-Conditions contracts to support legal experts in identifying and assessing problematic issues.
Outcome: The proposed annotation scheme achieves accuracies ranging from .79 to .95 on validation tasks.
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English (2022.acl-long)

Copied to clipboard

Challenge: Laws and their interpretations, legal arguments and agreements are typically expressed in writing.
Approach: They propose a benchmark to evaluate model performance across legal NLU tasks . they also evaluate several generic and legal-oriented models .
Outcome: The proposed model performs better across multiple tasks than previous models.
JurisBench: A Deep Benchmark for Assessing Large Language Models in Professional Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing legal benchmarks evaluate isolated tasks or exam-style questions, failing to capture the procedural interdependencies and adjudicative rigor inherent in professional practice.
Approach: They propose a vertical, depth-oriented, domain-specific benchmark to evaluate Large Language Models (LLMs) in Chinese civil litigation.
Outcome: The proposed benchmarks show that large language models exhibit an "illusion of competence" the results highlight a critical gap between fluent linguistic output and judicial reliability .
ACORD: An Expert-Annotated Retrieval Dataset for Legal Contract Drafting (2025.acl-long)

Copied to clipboard

Challenge: Contract clause retrieval is critical to contract drafting because of its high quality and complexity.
Approach: They propose the first expert-annotated benchmark specifically designed for contract clause retrieval . ACORD focuses on complex contract clauses such as Limitation of Liability, Indemnification, Change of Control .
Outcome: The atticus clause retrieval dataset shows promising results but needs improvement . the benchmark can be used as an IR benchmark for the NLP community .
OpenExempt: A Diagnostic Benchmark for Legal Reasoning and a Framework for Creating Custom Benchmarks on Demand (2026.findings-acl)

Copied to clipboard

Challenge: Reasoning benchmarks are expensive to build and ill suited for isolating specific failure modes.
Approach: They propose a framework and benchmark for diagnostic evaluation of legal reasoning that uses symbolic representations of U.S. Bankruptcy Code statutes to generate large space of reasoning tasks and their machine-computable solutions on demand.
Outcome: The proposed framework and benchmark provides diagnostic insights into the competencies and failure modes of language models.
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Frontier models often lack a view of performance on open-ended, economically consequential tasks in high-stakes professional domains where practical returns matter most.
Approach: They introduce a professional reasoning benchmark that recruits 182 qualified professionals to contribute questions inspired by their workflows.
Outcome: The proposed model outperforms other models in 114 countries and 47 US jurisdictions on hard subsets.
PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning.
Approach: They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios.
Outcome: The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics.
Pub-LawBench: Public-Oriented Benchmarking for LegalAI (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on legal professionals, not legal professionals.
Approach: They propose a public-oriented LegalAI benchmark grounded in legal functionalism and genre analysis to address this gap.
Outcome: The proposed model evaluates 17 large language models on Pub-LawBench using simple prompts and Chain-of-Thought under a vanilla inference setting.
What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions (2023.emnlp-main)

Copied to clipboard

Challenge: Existing systems that generate section-wise summaries of contracts can be tedious due to length and complexity of legalese.
Approach: They propose a task of party-specific extractive summarization for legal contracts . they train a pairwise importance ranker and propose incorporating domain-specific notions of importance .
Outcome: The proposed system generates a party-specific contract summary using a dataset of lease agreements and lease agreements.
LJPCheck: Functional Tests for Legal Judgment Prediction (2024.findings-acl)

Copied to clipboard

Challenge: Existing LJP models fail to evaluate specific aspects of their performance, such as legal fairness and judicial fairness.
Approach: They propose a suite of functional tests for LJP models to comprehend LJp models’ behaviors and offer diagnostic insights.
Outcome: Extensive tests reveal weaknesses in LJP models and provide diagnostic insights.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations