A Contract Corpus for Recognizing Rights and Obligations (2020.lrec-1)

Copied to clipboard

Challenge: Understanding the content of a contract is often difficult and costly, especially if the contract is long and complex.
Approach: They describe how they built an annotated corpus of contract documents that can be used to recognize rights and obligations.
Outcome: The proposed system can recognize parties' rights and obligations based on 46 English contracts and 25 Japanese contracts drafted by lawyers.

Similar Papers

ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts (2021.findings-emnlp)

Copied to clipboard

Challenge: Contract review is a time-consuming procedure that costs companies millions of dollars each year . linguistic characteristics of contracts, such as negations by exceptions, contribute to the difficulty of this task .
Approach: They propose a document-level natural language inference (NLI) task for contracts . they annotate and release the largest corpus to date consisting of 607 annotated contracts a linguistically rich system is proposed .
Outcome: The proposed system is based on a contract review task that includes 607 annotated contracts.
Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts (2024.lrec-main)

Copied to clipboard

Challenge: Using Large Language Models (LLMs) as foundational models, we propose a new annotation scheme to classify different types of clauses in Terms-and-Conditions contracts.
Approach: They propose to use a new annotation scheme to classify clauses in Terms-and-Conditions contracts to support legal experts in identifying and assessing problematic issues.
Outcome: The proposed annotation scheme achieves accuracies ranging from .79 to .95 on validation tasks.
What to Read in a Contract? Party-Specific Summarization of Legal Obligations, Entitlements, and Prohibitions (2023.emnlp-main)

Copied to clipboard

Challenge: Existing systems that generate section-wise summaries of contracts can be tedious due to length and complexity of legalese.
Approach: They propose a task of party-specific extractive summarization for legal contracts . they train a pairwise importance ranker and propose incorporating domain-specific notions of importance .
Outcome: The proposed system generates a party-specific contract summary using a dataset of lease agreements and lease agreements.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
LEDGAR: A Large-Scale Multi-label Corpus for Text Classification of Legal Provisions in Contracts (2020.lrec-1)

Copied to clipboard

Challenge: Contractual provisions are a primary research target in law studies as they constitute the legal essence of a contract.
Approach: They propose to use LEDGAR to construct a multilabel corpus of legal provisions in contracts that is crawled and scraped from the public domain.
Outcome: The proposed corpus is the first freely available corpus of its kind.
Corpus for Automatic Structuring of Legal Documents (2022.lrec-1)

Copied to clipboard

Challenge: In populous countries, pending legal cases are growing exponentially.
Approach: They propose a corpus of legal judgment documents in English that is annotated with a label coming from a list of pre-defined rhetorical roles.
Outcome: The proposed corpus of legal judgment documents is annotated with a label coming from a list of pre-defined rhetorical roles.
A Legal Perspective on Training Models for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: a significant concern in processing natural language data is the unclear legal status of the input and output data/resources.
Approach: They examine which legal rules apply at relevant steps and how they affect the legal status of the results.
Outcome: The proposed model training process is based on three scenarios . the analysis focuses on which legal rules apply and how they affect the legal status of the results .
Contract Discovery: Dataset and a Few-Shot Semantic Retrieval Challenge with Competitive Baselines (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for detecting text fragments are not suitable for contract discovery, since it requires manual definition of a few examples, followed by conventional information.
Approach: They propose a task where legal clauses are extracted from documents, given a few examples of similar clauses from other legal acts.
Outcome: The proposed task differs substantially from conventional NLI and shared tasks on legal information extraction.
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts.
Approach: They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent .
Outcome: The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora .
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations