Challenge: Existing models for identifying/retrieving relevant statutes and prior cases/precedents are inherently related, e.g., similar cases tend to cite similar statutes due to similar factual situation.
Approach: They propose a corpus that provides a common testbed for developing models that exploit the dependence between the two tasks.
Outcome: The proposed corpus exploits the dependence between the two retrieval tasks and provides a baseline model for the two tasks.

Similar Papers

U-CREAT: Unsupervised Case Retrieval using Events extrAcTion (2023.acl-long)

Copied to clipboard

Challenge: Prior Case Retrieval (PCR) is about automatically citing relevant prior legal cases in a given query case.
Approach: They propose a new benchmark for prior case retrieval based on a legal query case . they propose an unsupervised retrieval method-based pipeline U-CREAT .
Outcome: The proposed method significantly improves performance and makes retrieval faster compared to BM25.
Legal Case Retrieval: A Survey of the State of the Art (2024.acl-long)

Copied to clipboard

Challenge: Recent years have seen increasing attention on Legal Case Retrieval (LCR) this task involves retrieving cases from a legal database of historical cases that are similar to a given query case.
Approach: They present a survey of the major milestones made in legal case retrieval research . they seek to understand the datasets and recent neural models and their performances .
Outcome: The proposed task is based on a dataset of historical cases similar to a given query case.
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Legal systems worldwide struggle with exponentially growing legal cases in various courts.
Approach: They propose a benchmark for Indian legal text understanding and reasoning task that includes domain-specific tasks that address different aspects of the legal system.
Outcome: The proposed benchmark for Indian legal text understanding and reasoning aims to address the gap between models and the ground truth.
ECtHR-PCR: A Dataset for Precedent Understanding and Prior Case Retrieval in the European Court of Human Rights (2024.lrec-main)

Copied to clipboard

Challenge: Prior case retrieval datasets do not simulate a realistic setting because they use complete case documents while only masking references to prior cases.
Approach: They propose a prior case retrieval dataset based on judgements from the European Court of Human Rights which explicitly separate facts from arguments and exhibit precedential practices.
Outcome: The proposed datasets do not simulate a realistic setting and expose queries to spurious patterns left behind by citation masks, potentially short-circuiting a comprehensive understanding of case facts and legal principles.
Finding the Law: Enhancing Statutory Article Retrieval via Graph Neural Networks (2023.eacl-main)

Copied to clipboard

Challenge: Statutory article retrieval (SAR) is a promising application of legal text processing.
Approach: They propose a graph-augmented dense statute retriever model that incorporates the structure of legislation via a neural network to improve density retrieval performance.
Outcome: The proposed model outperforms baselines on a real-world expert-annotated dataset.
LeCoPCR: Legal Concept-guided Prior Case Retrieval for European Court of Human Rights cases (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches overlook the underlying semantic intent in determining relevance with respect to a query case.
Approach: They propose a method that generates intents in the form of legal concepts from a query case facts and then augments the query with these concepts to enhance models understanding of semantic intent.
Outcome: The proposed approach generates intents in the form of legal concepts and augments the query with these concepts to enhance models understanding of semantic intent that dictates relavance.
ILSIC: Corpora for Identifying Indian Legal Statutes from Queries by Laymen (2026.findings-eacl)

Copied to clipboard

Challenge: Existing studies have focused on the use of court judgments as input for legal Statute Identification (LSI) however, there is little research to explore the differences between court and laypeople data for LSI.
Approach: They create a corpus of laypeople queries covering 500+ statutes from Indian law . they use court case judgements to compare between the two datasets .
Outcome: The proposed corpus of laypeople queries covers 500+ statutes from Indian law . the results show that models trained on court judgements are ineffective .
LePaRD: A Large-Scale Dataset of Judicial Citations to Precedent (2024.acl-long)

Copied to clipboard

Challenge: Legal passage retrieval is a practice-oriented task that seeks to predict relevant passages from precedential court decisions given the context of a legal argument.
Approach: They present a dataset which aims to facilitate work on legal passage retrieval . they extensively evaluate various approaches and find classification-based retrieval works best .
Outcome: The proposed dataset aims to facilitate work on legal passage retrieval . it shows that classification-based retrieval seems to work best .
CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation (2025.findings-naacl)

Copied to clipboard

Challenge: a dataset of case law is used to train and evaluate models for writing legal analyses . current approaches struggle to find relevant cases and generate legal analyses, authors say .
Approach: They build a dataset of case law to support information retrieval and retrieval-augmented generation.
Outcome: The proposed dataset supports two important backbone tasks: retrieval (IR) and retrieval-augmented generation (RAG).
CaseEncoder: A Knowledge-enhanced Pre-trained Model for Legal Case Encoding (2023.emnlp-main)

Copied to clipboard

Challenge: Existing legal-oriented PLMs rely on replacing general domain training data with legal data or extending the input length to fit the long-length characteristic of legal data.
Approach: They propose a legal document encoder that leverages fine-grained legal knowledge in both the data sampling and pre-training phases.
Outcome: The proposed model outperforms existing general domain pre-training models and legal-specific pre-trainers on multiple benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations