AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer Contracts (2024.acl-long)
Copied to clipboard
| Challenge: | Annotated contracts are laborious task performed by companies, law firms, NGOs and the scientific community. |
| Approach: | They present a corpus of 3,764 clauses from German consumer contracts annotated by legal experts with a clause in the contract. |
| Outcome: | The proposed framework outperforms openly available models in detecting potentially void clauses. |
Similar Papers
A Dataset of German Legal Documents for Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | a dataset developed for Named Entity Recognition in German federal court decisions is available under a CC-BY 4.0 license. |
| Approach: | They describe a dataset developed for Named Entity Recognition in German federal court decisions. |
| Outcome: | The proposed dataset was developed for training an NER service for German legal documents in the EU project Lynx. |
Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts (2024.lrec-main)
Copied to clipboard
| Challenge: | Using Large Language Models (LLMs) as foundational models, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts. |
| Approach: | They propose to use a new annotation scheme to classify clauses in Terms-and-Conditions contracts to support legal experts in identifying and assessing problematic issues. |
| Outcome: | The proposed annotation scheme achieves accuracies ranging from .79 to .95 on validation tasks. |
Answering legal questions from laymen in German civil law system (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing studies have focused on questions asked by experts, such as lawyers or legal scholars. |
| Approach: | They use a dataset to analyze laymen's legal questions paired with answers from lawyers and grounded to concrete law book paragraphs to find out what limitations exist. |
| Outcome: | The proposed system could help laymen in real situations without understanding law . the proposed system is based on 21k laymen’s legal questions paired with answers from lawyers and grounded to concrete law book paragraphs. |
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English (2022.acl-long)
Copied to clipboard
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, Nikolaos Aletras
| Challenge: | Laws and their interpretations, legal arguments and agreements are typically expressed in writing. |
| Approach: | They propose a benchmark to evaluate model performance across legal NLU tasks . they also evaluate several generic and legal-oriented models . |
| Outcome: | The proposed model performs better across multiple tasks than previous models. |
NESTLE: a No-Code Tool for Statistical Analysis of Legal Corpus (2024.eacl-demo)
Copied to clipboard
| Challenge: | a comprehensive statistical analysis of legal corpus requires specialized tools or programming skills. |
| Approach: | They propose a no-code tool for large-scale statistical analysis of legal corpus . NESTLE can extract any type of information that has not been predefined in the IE system . |
| Outcome: | The proposed tool can perform comparable to LexGLUE on 15 Korean precedent IE tasks and 3 legal text classification tasks. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation (2025.findings-naacl)
Copied to clipboard
Abe Bohan Hou, Orion Weller, Guanghui Qin, Eugene Yang, Dawn Lawrie, Nils Holzenberger, Andrew Blair-Stanek, Benjamin Van Durme
| Challenge: | a dataset of case law is used to train and evaluate models for writing legal analyses . current approaches struggle to find relevant cases and generate legal analyses, authors say . |
| Approach: | They build a dataset of case law to support information retrieval and retrieval-augmented generation. |
| Outcome: | The proposed dataset supports two important backbone tasks: retrieval (IR) and retrieval-augmented generation (RAG). |
LEDGAR: A Large-Scale Multi-label Corpus for Text Classification of Legal Provisions in Contracts (2020.lrec-1)
Copied to clipboard
| Challenge: | Contractual provisions are a primary research target in law studies as they constitute the legal essence of a contract. |
| Approach: | They propose to use LEDGAR to construct a multilabel corpus of legal provisions in contracts that is crawled and scraped from the public domain. |
| Outcome: | The proposed corpus is the first freely available corpus of its kind. |
Information Extraction from Legal Wills: How Well Does GPT-4 Do? (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Using information extraction from legal wills is an important application of artificial intelligence (AI) |
| Approach: | They propose a manually annotated dataset for Information Extraction (IE) from legal wills . they also use it to evaluate the performance of large language models (LLMs) |
| Outcome: | The proposed dataset can be used to evaluate large language models on IE from legal wills . it shows that the model performs reasonably well, but inconsistent outputs and overgeneralization are observed . |
Annotation and Automatic Classification of Aspectual Categories (P19-1)
Copied to clipboard
| Challenge: | Annotated resource for aspectual classification of German verb tokens in context. |
| Approach: | They present a resource for aspectual classification of German verb tokens in their clausal context. |
| Outcome: | The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications. |