Challenge: a dataset of 20K rulings from the Swiss Federal Supreme Court is lacking in legal headnotes due to the high cost of manual annotation.
Approach: They propose a dataset that contains 20K rulings from the Swiss Federal Supreme Court . they fine-tune open models and compare them to larger general-purpose and reasoning-tunned LLMs .
Outcome: The proposed dataset contains 20K rulings from the Swiss Federal Supreme Court with headnotes in German, French, and Italian.

Similar Papers

SwiLTra-Bench: The Swiss Legal Translation Benchmark (2025.acl-long)

Copied to clipboard

Challenge: In Switzerland legal translation relies on legal experts who must be both legal experts and skilled translators—creating bottlenecks and impacting effective access to justice.
Approach: They propose a multilingual benchmarking system that evaluates Swiss legal translation systems based on 180K aligned Swiss legal translator pairs . they show frontier models achieve superior translation performance across all document types while specialized translation systems excel specifically in laws but under-perform in headnotes.
Outcome: The proposed model outperforms specialized models in laws but underperform in headnotes.
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence (2025.acl-short)

Copied to clipboard

Challenge: Existing approaches to evaluating the importance of legal cases are manual and resource-intensive.
Approach: They propose a dataset that uses two-tier labels to evaluate case criticality . they use the LD-Label to identify cases published as Leading Decisions and the Citation-L Label to rank cases by their citation frequency and recency.
Outcome: The Criticality Prediction dataset outperforms existing approaches to evaluate case criticality . the proposed model outperformed the existing models in a zero-shot setting .
LexAbSumm: Aspect-based Summarization of Legal Decisions (2024.lrec-main)

Copied to clipboard

Challenge: LexAbSumm is a dataset designed for aspect-based summarization of legal documents . it is based on a set of ECtHR fact sheets, and is available for download.
Approach: They propose a dataset designed for aspect-based summarization of legal case decisions . they evaluate abstractive summarizing models tailored for longer documents .
Outcome: The proposed dataset is designed for aspect-based summarization of legal cases . it reveals a challenge in conditioning models to produce aspect-specific summaries .
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Using Swiss Judgement Prediction, we evaluate the explainability of state-of-the-art monolingual and multilingual LJP models.
Approach: They propose an occlusion-based approach to evaluate the explainability performance of legal judgement prediction models using Swiss Judgement Prediction, the only available multilingual LJP dataset.
Outcome: The proposed framework allows us to quantify the influence of lower court information on model predictions, exposing current models’ biases.
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions (2025.findings-naacl)

Copied to clipboard

Challenge: CaseSumm is a dataset for long-context summarization in the legal domain . human groundtruth summaries are often not available for legal summarizing .
Approach: They propose a dataset for long-context summarization that includes SCOTUS opinions and their official summaries.
Outcome: The proposed dataset is the largest open legal case summarization dataset . it outperforms larger models on automatic metrics and human evaluation .
EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain (2022.emnlp-main)

Copied to clipboard

Challenge: Existing summarization datasets focus on overly exposed domains and are primarily monolingual with few multilingual datasets.
Approach: They propose a new summarization dataset based on manually curated document summaries from the European Union law platform EUR-Lex.
Outcome: The proposed dataset is based on document summaries of legal acts from the European Union law platform (EUR-Lex).
Modeling Legal Reasoning: LM Annotation at the Edge of Human Agreement (2023.emnlp-main)

Copied to clipboard

Challenge: Existing research examines simple classification tasks, but ability of LMs to classify on complex tasks is less well understood.
Approach: They analyze a Supreme Court opinion annotated by a team of domain experts . they find generative models perform poorly when given instructions equal to human annotators .
Outcome: The proposed model performs poorly when given instructions equal to instructions given to human annotations . strongest results derive from fine-tuning models on the annotated dataset .
CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis (2025.findings-emnlp)

Copied to clipboard

Challenge: a new benchmark is designed to evaluate LLMs on Chinese legal knowledge and its application in reasoning . general pre-training that ingests legal texts without specialized focus compromises reliability of LLM responses . achieving trustworthy legal reasoning in LLM requires a robust synergy of accurate knowledge retrieval and strong general reasoning capabilities.
Approach: They propose a benchmark specifically engineered to evaluate LLMs on Chinese legal knowledge and its application in reasoning.
Outcome: The proposed benchmark evaluates LLMs on Chinese legal knowledge and its application in reasoning.
Complex Labelling and Similarity Prediction in Legal Texts: Automatic Analysis of France’s Court of Cassation Rulings (2022.lrec-1)

Copied to clipboard

Challenge: Detecting divergences in the applications of the law is an important task . divergencies can occur at three levels: within the Cour de Cassation, between trial courts and, more rarely, between a trial court and the Cour of Cassion.
Approach: They propose to provide automatic tools to facilitate the search for similar rulings . they provide automatic keyword sequence generation models and predict keyword sequences based on available texts .
Outcome: The proposed tools improve correlations between the obtained similarities and human judgments of similarity.
Extractive Summarization of Legal Decisions using Multi-task Learning and Maximal Marginal Relevance (2022.findings-emnlp)

Copied to clipboard

Challenge: Summarizing legal decisions requires the expertise of law practitioners, which is time- and cost-intensive.
Approach: They propose methods for extracting summarized legal decisions using limited expert annotated data.
Outcome: The proposed models achieve ROUGE scores vis-à-vis expert extracted summaries that match inter-annotator comparisons.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations