Challenge: Existing studies focus on case-to-case retrieval using lengthy queries, which does not match real-world scenarios.
Approach: They propose a method to construct query-candidate pairs and build the largest LCR dataset to date, LEAD.
Outcome: Experimental results show that the method can provide ample training signals for LCR models.

Similar Papers

LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on legal case retrieval have limited results . limited representations and legally irrelevant matches are often used .
Approach: They propose a large-scale Korean LCR benchmark and a retrieval model that performs legal element reasoning over the query case.
Outcome: a new model outperforms baseline models on a Korean LCR benchmark . it performs state-of-the-art on 411 diverse crime types in queries over 1.2M candidate cases . previous studies have shown that the model can generalize to out-of domain cases if it is trained on in-domain data .
Legal Case Retrieval: A Survey of the State of the Art (2024.acl-long)

Copied to clipboard

Challenge: Recent years have seen increasing attention on Legal Case Retrieval (LCR) this task involves retrieving cases from a legal database of historical cases that are similar to a given query case.
Approach: They present a survey of the major milestones made in legal case retrieval research . they seek to understand the datasets and recent neural models and their performances .
Outcome: The proposed task is based on a dataset of historical cases similar to a given query case.
GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing dense retrieval methods neglect the explicit legal logic that underpins legal relevance.
Approach: They propose a framework that reformulates retrieval as an inference process over latent legal variables.
Outcome: GLIER outperforms strong baselines like SAILER and KELLER in a legal case-based retrieval task . the framework exhibits exceptional data efficiency even when trained with only 10% of the data .
LexCLiPR: Cross-Lingual Paragraph Retrieval from Legal Judgments (2025.acl-long)

Copied to clipboard

Challenge: Existing work on IR focus on retrieving entire cases rather than precise, paragraph-level information.
Approach: They propose a cross-lingual dataset for paragraph-level retrieval from ECtHR judgments . they evaluate retrieval models in a zero-shot setting and use multilingual case law guides .
Outcome: The proposed model excels in cross-lingual retrieval, while siamese architectures are better suited for monolingual tasks.
ECtHR-PCR: A Dataset for Precedent Understanding and Prior Case Retrieval in the European Court of Human Rights (2024.lrec-main)

Copied to clipboard

Challenge: Prior case retrieval datasets do not simulate a realistic setting because they use complete case documents while only masking references to prior cases.
Approach: They propose a prior case retrieval dataset based on judgements from the European Court of Human Rights which explicitly separate facts from arguments and exhibit precedential practices.
Outcome: The proposed datasets do not simulate a realistic setting and expose queries to spurious patterns left behind by citation masks, potentially short-circuiting a comprehensive understanding of case facts and legal principles.
CDD: A Large Scale Dataset for Legal Intelligence Research (2023.emnlp-industry)

Copied to clipboard

Challenge: Recent research has focused on predicting crimes, predicting outcomes of judicial debates, and extracting information from legal documents.
Approach: They propose to use a large-size Court Debate Dataset to analyze court debates . they invite experienced judges to design appropriate labels for data records .
Outcome: The proposed dataset includes 30,481 court cases, totaling 1,144,425 utterances.
CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation (2025.findings-naacl)

Copied to clipboard

Challenge: a dataset of case law is used to train and evaluate models for writing legal analyses . current approaches struggle to find relevant cases and generate legal analyses, authors say .
Approach: They build a dataset of case law to support information retrieval and retrieval-augmented generation.
Outcome: The proposed dataset supports two important backbone tasks: retrieval (IR) and retrieval-augmented generation (RAG).
Logic Rules as Explanations for Legal Case Retrieval (2024.lrec-main)

Copied to clipboard

Challenge: Recent efforts to learn explainable legal case retrieval models fail to provide faithful and interpretable explanations for legal cases.
Approach: They propose a framework that uses logic rules to explain legal case retrieval results . they extend benchmarks of LeCaRD and ELAM with manually annotated logic rules .
Outcome: The proposed framework is able to provide faithful explanations for legal case retrieval.
CaseEncoder: A Knowledge-enhanced Pre-trained Model for Legal Case Encoding (2023.emnlp-main)

Copied to clipboard

Challenge: Existing legal-oriented PLMs rely on replacing general domain training data with legal data or extending the input length to fit the long-length characteristic of legal data.
Approach: They propose a legal document encoder that leverages fine-grained legal knowledge in both the data sampling and pre-training phases.
Outcome: The proposed model outperforms existing general domain pre-training models and legal-specific pre-trainers on multiple benchmarks.
LeCoPCR: Legal Concept-guided Prior Case Retrieval for European Court of Human Rights cases (2025.findings-naacl)

Copied to clipboard

Challenge: Existing approaches overlook the underlying semantic intent in determining relevance with respect to a query case.
Approach: They propose a method that generates intents in the form of legal concepts from a query case facts and then augments the query with these concepts to enhance models understanding of semantic intent.
Outcome: The proposed approach generates intents in the form of legal concepts and augments the query with these concepts to enhance models understanding of semantic intent that dictates relavance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations