Challenge: Existing retrieval methods struggle to achieve ideal results, a study finds . existing large language models lack prior knowledge of the content of superior legal articles .
Approach: They propose to use a Chinese superior legal article retrieval dataset to find relevant articles with higher legal effectiveness.
Outcome: The proposed dataset shows that existing retrieval methods struggle to achieve ideal results.

Similar Papers

A Statutory Article Retrieval Dataset in French (2022.acl-long)

Copied to clipboard

Challenge: Statutory article retrieval is the task of automatically retrieving law articles relevant to a legal question.
Approach: They propose to use a Belgian Statutory Article Retrieval Dataset to test various retrieval approaches including lexical and dense architectures to achieve a 74.8% R@100.
Outcome: The proposed dataset outperforms existing systems in both zero-shot and supervised setups.
An Element is Worth a Thousand Words: Enhancing Legal Case Retrieval by Incorporating Legal Elements (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for legal case retrieval lack the definition of relevance for legal cases . however, the definition goes beyond the common semantic relevance of ad-hoc retrieval.
Approach: They propose a legal element dataset that incorporates legal elements into a semi-automatic method . they propose two models to enhance legal search using legal elements .
Outcome: The proposed models outperform existing methods in enhancing legal search using legal elements.
LePaRD: A Large-Scale Dataset of Judicial Citations to Precedent (2024.acl-long)

Copied to clipboard

Challenge: Legal passage retrieval is a practice-oriented task that seeks to predict relevant passages from precedential court decisions given the context of a legal argument.
Approach: They present a dataset which aims to facilitate work on legal passage retrieval . they extensively evaluate various approaches and find classification-based retrieval works best .
Outcome: The proposed dataset aims to facilitate work on legal passage retrieval . it shows that classification-based retrieval seems to work best .
Legal Case Retrieval: A Survey of the State of the Art (2024.acl-long)

Copied to clipboard

Challenge: Recent years have seen increasing attention on Legal Case Retrieval (LCR) this task involves retrieving cases from a legal database of historical cases that are similar to a given query case.
Approach: They present a survey of the major milestones made in legal case retrieval research . they seek to understand the datasets and recent neural models and their performances .
Outcome: The proposed task is based on a dataset of historical cases similar to a given query case.
CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation (2025.findings-naacl)

Copied to clipboard

Challenge: a dataset of case law is used to train and evaluate models for writing legal analyses . current approaches struggle to find relevant cases and generate legal analyses, authors say .
Approach: They build a dataset of case law to support information retrieval and retrieval-augmented generation.
Outcome: The proposed dataset supports two important backbone tasks: retrieval (IR) and retrieval-augmented generation (RAG).
LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies on legal case retrieval have limited results . limited representations and legally irrelevant matches are often used .
Approach: They propose a large-scale Korean LCR benchmark and a retrieval model that performs legal element reasoning over the query case.
Outcome: a new model outperforms baseline models on a Korean LCR benchmark . it performs state-of-the-art on 411 diverse crime types in queries over 1.2M candidate cases . previous studies have shown that the model can generalize to out-of domain cases if it is trained on in-domain data .
GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing dense retrieval methods neglect the explicit legal logic that underpins legal relevance.
Approach: They propose a framework that reformulates retrieval as an inference process over latent legal variables.
Outcome: GLIER outperforms strong baselines like SAILER and KELLER in a legal case-based retrieval task . the framework exhibits exceptional data efficiency even when trained with only 10% of the data .
STARD: A Chinese Statute Retrieval Dataset Derived from Real-life Queries by Non-professionals (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing statute retrieval benchmarks emphasize formal and professional queries from sources like bar exams and legal case documents . existing retrieval approaches that lack domain-specific knowledge may struggle to capture the meanings of specialized terms accurately.
Approach: They propose a dataset that captures the complexity and diversity of real queries from the general public.
Outcome: The proposed dataset captures the complexity and diversity of real queries from the general public.
CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: a new benchmark is designed to evaluate LLMs on Chinese legal knowledge and its application in reasoning . general pre-training that ingests legal texts without specialized focus compromises reliability of LLM responses . achieving trustworthy legal reasoning in LLM requires a robust synergy of accurate knowledge retrieval and strong general reasoning capabilities.
Approach: They propose a benchmark specifically engineered to evaluate LLMs on Chinese legal knowledge and its application in reasoning.
Outcome: The proposed benchmark evaluates LLMs on Chinese legal knowledge and its application in reasoning.
Improving Vietnamese-English Cross-Lingual Retrieval for Legal and General Domains (2025.naacl-short)

Copied to clipboard

Challenge: Existing document retrieval systems focus on a single language, targeting resource-rich languages like English or Chinese.
Approach: They propose auxiliary loss function and symmetrical training strategy for cross-lingual retrieval between Vietnamese and English . they propose a dataset that covers the general domain and extends to the legal field .
Outcome: The proposed dataset significantly improves state-of-the-art models on cross-lingual retrieval tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations