Challenge: In general, instruction tuning is important for direct user interaction, but the legal domain is underrepresented in typical instruction datasets.
Approach: They aggregate 58 annotated legal datasets and write instructions for each to create LawInstruct.
Outcome: The proposed model improves on LegalBench across all model sizes, but no drop in MMLU.

Similar Papers

Modeling Legal Reasoning: LM Annotation at the Edge of Human Agreement (2023.emnlp-main)

Copied to clipboard

Challenge: Existing research examines simple classification tasks, but ability of LMs to classify on complex tasks is less well understood.
Approach: They analyze a Supreme Court opinion annotated by a team of domain experts . they find generative models perform poorly when given instructions equal to human annotators .
Outcome: The proposed model performs poorly when given instructions equal to instructions given to human annotations . strongest results derive from fine-tuning models on the annotated dataset .
LEMUR: A Corpus for Robust Fine-Tuning of Multilingual Law Embedding Models for Retrieval (2026.eacl-srw)

Copied to clipboard

Challenge: Existing large language models are not designed for semantic retrieval and PDF-based legislative sources introduce substantial noise due to imperfect text extraction.
Approach: They propose a large-scale multilingual corpus of EU environmental legislation constructed from 24,953 official EUR-Lex PDF documents covering 25 languages.
Outcome: The proposed model improves Top-k retrieval accuracy in monolingual and bilingual settings . it also improves accuracy in low- and high-resource languages .
DS2-Instruct: Domain-Specific Data Synthesis for Large Language Models Instruction Tuning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing data synthesis methods focus on general-purpose tasks and fail to capture domain-specific terminology and reasoning patterns.
Approach: They propose a framework that generates domain-specific instruction datasets without human supervision by pairing task-informed keywords with different cognitive levels from Bloom’s Taxonomy.
Outcome: The proposed framework generates domain-specific instruction datasets without human supervision and achieves significant improvements over existing methods.
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh (2025.acl-long)

Copied to clipboard

Challenge: Instruction tuning in low-resource languages remains underexplored due to limited text data, particularly in government and cultural domains.
Approach: They propose to open-source a large-scale instruction-following dataset covering key institutional and cultural knowledge relevant to Kazakhstan.
Outcome: The proposed dataset improves LLMs’ understanding of procedural, legal, and structural governance topics.
A Legal Perspective on Training Models for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: a significant concern in processing natural language data is the unclear legal status of the input and output data/resources.
Approach: They examine which legal rules apply at relevant steps and how they affect the legal status of the results.
Outcome: The proposed model training process is based on three scenarios . the analysis focuses on which legal rules apply and how they affect the legal status of the results .
LEGAL-BERT: The Muppets straight out of Law School (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing guidelines for pre-training and fine-tuning do not always generalize well in the legal domain.
Approach: They propose to use BERT out of the box, adapt it by additional pre-training on domain-specific corpora, and pre-train it from scratch on domains.
Outcome: The proposed strategies are: use the original BERT out of the box, adapt it by additional pre-training on domain-specific corpora, and pre-train it from scratch on domain specific corpors.
Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference (2025.emnlp-main)

Copied to clipboard

Challenge: Reasoning-focused large language models (LLMs) are rapidly evolving across various domains, yet their capabilities in handling complex legal problems remain underexplored.
Approach: They propose a large language model tailored for legal reasoning with a 7-billion parameter scale and a two-stage training strategy combining Supervised Fine-Tuning and Reinforcement Learning.
Outcome: The proposed model outperforms all models of similar scale on authoritative benchmarks and outperformed Qwen-2.5-7B-Instruct (46.6%) by an average margin of 6.6%.
LawBench: Benchmarking Legal Knowledge of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: LegalBench evaluated 20 LLMs in 162 legal tasks in 20 countries and jurisdictions.
Approach: They present a comprehensive evaluation of 21 popular Large Language Models and the first comparative analysis of the empirical results.
Outcome: The proposed benchmarks are based on the Bloom’s cognitive taxonomy and are compared to 21 popular LLMs.
ArgInstruct: Specialized Instruction Fine-Tuning for Computational Argumentation (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have been trained to follow instructions for many NLP tasks, including several tasks from computational argumentation (CA), the computational analysis and synthesis of natural language arguments.
Approach: They propose a specialized instruction fine-tuning for the domain of computational argumentation (CA) they synthesized 52k CA-related instructions and used them to train a CA-specialized instruction-following LLM.
Outcome: The proposed benchmarks show that the LLMs can tackle unseen and seen tasks while maintaining generalization capabilities.
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English (2022.acl-long)

Copied to clipboard

Challenge: Laws and their interpretations, legal arguments and agreements are typically expressed in writing.
Approach: They propose a benchmark to evaluate model performance across legal NLU tasks . they also evaluate several generic and legal-oriented models .
Outcome: The proposed model performs better across multiple tasks than previous models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations