| Challenge: | formalizing legal sources is an important challenge, but the generation of a formal representation from legal texts has been less considered and requires considerable expertise. |
| Approach: | They propose to experiment with annotations and the annotation process to improve uniformity and efficiency of legal annotation. |
| Outcome: | The proposed method improves the richness and efficiency of legal annotations. |
Similar Papers
Towards Automated Extraction of Business Constraints from Unstructured Regulatory Text (C18-2)
Copied to clipboard
| Challenge: | a system for machine-driven annotations of legal documents is currently undergoing user trials within our organization. |
| Approach: | a system for machine-driven annotations of legal documents is presented . the system is currently undergoing user trials within our organization. |
| Outcome: | the proposed system is currently undergoing user trials within our organization. |
A Legal Perspective on Training Models for Natural Language Processing (L18-1)
Copied to clipboard
| Challenge: | a significant concern in processing natural language data is the unclear legal status of the input and output data/resources. |
| Approach: | They examine which legal rules apply at relevant steps and how they affect the legal status of the results. |
| Outcome: | The proposed model training process is based on three scenarios . the analysis focuses on which legal rules apply and how they affect the legal status of the results . |
LegalViz: Legal Text Visualization by Text To Diagram Generation (2025.naacl-long)
Copied to clipboard
| Challenge: | Graphviz provides diagrams for legal documents that are easy to understand and understand . a novel dataset of 23 languages and 7,010 cases of legal document and visualization pairs is proposed . |
| Approach: | They propose a dataset of legal diagrams using DOT graph description language of Graphviz. |
| Outcome: | The proposed dataset outperforms existing models including GPTs in 23 languages and 7,010 cases of legal document and visualization pairs. |
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English (2022.acl-long)
Copied to clipboard
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, Nikolaos Aletras
| Challenge: | Laws and their interpretations, legal arguments and agreements are typically expressed in writing. |
| Approach: | They propose a benchmark to evaluate model performance across legal NLU tasks . they also evaluate several generic and legal-oriented models . |
| Outcome: | The proposed model performs better across multiple tasks than previous models. |
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text (2024.eacl-long)
Copied to clipboard
Dor Bernsohn, Gil Semo, Yaron Vazana, Gila Hayat, Ben Hagag, Joel Niklaus, Rohit Saha, Kyryl Truskovskyi
| Challenge: | a recent study focused on detecting legal violations within unstructured textual data . a similar study focused only on associating violations with potentially affected individuals . |
| Approach: | They constructed two datasets using Large Language Models (LLMs) they publicize the results to advance legal natural language processing research . |
| Outcome: | The proposed datasets and the code used for the experiments have been released to advance legal natural language processing (NLP) |
Corpus for Automatic Structuring of Legal Documents (2022.lrec-1)
Copied to clipboard
Prathamesh Kalamkar, Aman Tiwari, Astha Agarwal, Saurabh Karn, Smita Gupta, Vivek Raghavan, Ashutosh Modi
| Challenge: | In populous countries, pending legal cases are growing exponentially. |
| Approach: | They propose a corpus of legal judgment documents in English that is annotated with a label coming from a list of pre-defined rhetorical roles. |
| Outcome: | The proposed corpus of legal judgment documents is annotated with a label coming from a list of pre-defined rhetorical roles. |
An Evaluation Framework for Legal Document Summarization (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing metrics for summarizing legal documents fail to evaluate intent in the original text. |
| Approach: | They propose an automated intent-based summarization metric which shows a better agreement with human evaluation as compared to other automated metrics like BLEU, ROUGE-L etc. |
| Outcome: | The proposed method shows that human evaluation is more accurate than other metrics. |
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation (2025.acl-long)
Copied to clipboard
| Challenge: | a novel framework for automated legal interpretation is proposed to alleviate the burden on legal experts. |
| Approach: | They propose a framework for automated legal interpretation that uses large language models to extract concept-related information and interpret legal concepts. |
| Outcome: | The proposed framework eliminates the need for legal experts to interpret legal concepts . it uses large language models to extract concept-related information and interpret legal concept interpretations . |
Collection and Annotation of the Romanian Legal Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, the corpus contains more than 140k documents representing the legislative body of Romania. |
| Approach: | They present a Romanian legislative corpus which is a valuable linguistic asset for machine translation systems. |
| Outcome: | The Romanian legislative corpus contains more than 140k documents representing the legislative body of Romania. |
Cross-lingual Annotation Projection in Legal Texts (2020.coling-main)
Copied to clipboard
| Challenge: | a new study examines annotation projection in text classification problems where source documents are published in multiple languages. |
| Approach: | They propose to use word embeddings and dynamic time warping to create an annotation corpus for text classification problems where source documents are published in multiple languages. |
| Outcome: | The proposed method is based on word embeddings and dynamic time warping . the aim is to train linguistic tools for the target language without experts . |