Challenge: In the context of the Indian judiciary, there is an additional complexity - Indian legal case judgments are mostly written in complex English due to historical reasons, but a significant portion of India's population lacks a strong command of the English language.
Approach: They propose to summarize Indian legal case judgments in English and Hindi by combining the summaries of 3,122 case judgment from Indian courts into one dataset.
Outcome: The proposed dataset compares the summarization methods with other datasets and shows that the proposed approaches perform better than previous approaches.

Similar Papers

IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Legal systems worldwide struggle with exponentially growing legal cases in various courts.
Approach: They propose a benchmark for Indian legal text understanding and reasoning task that includes domain-specific tasks that address different aspects of the legal system.
Outcome: The proposed benchmark for Indian legal text understanding and reasoning aims to address the gap between models and the ground truth.
PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for Indian languages are limited in terms of coverage and size.
Approach: They propose a multilingual and massively parallel summarization corpus focused on languages in India that provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
Outcome: The proposed dataset provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
A Workbench for Rapid Generation of Cross-Lingual Summaries (L18-1)

Copied to clipboard

Challenge: a tool for automating cross-lingual information access is needed in multilingual societies . current state of machine translation is not able to generate publishable articles from English .
Approach: They propose a web-based tool for human editing of cross-lingual summaries . it generates publishable summary in a number of Indian Languages for news articles originally published in english .
Outcome: The proposed tool can generate publishable summaries in multiple languages with minimal human effort and collect detailed logs on the process.
LexAbSumm: Aspect-based Summarization of Legal Decisions (2024.lrec-main)

Copied to clipboard

Challenge: LexAbSumm is a dataset designed for aspect-based summarization of legal documents . it is based on a set of ECtHR fact sheets, and is available for download.
Approach: They propose a dataset designed for aspect-based summarization of legal case decisions . they evaluate abstractive summarizing models tailored for longer documents .
Outcome: The proposed dataset is designed for aspect-based summarization of legal cases . it reveals a challenge in conditioning models to produce aspect-specific summaries .
EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain (2022.emnlp-main)

Copied to clipboard

Challenge: Existing summarization datasets focus on overly exposed domains and are primarily monolingual with few multilingual datasets.
Approach: They propose a new summarization dataset based on manually curated document summaries from the European Union law platform EUR-Lex.
Outcome: The proposed dataset is based on document summaries of legal acts from the European Union law platform (EUR-Lex).
HLDC: Hindi Legal Documents Corpus (2022.findings-acl)

Copied to clipboard

Challenge: Existing systems that process legal documents are lacking high-quality corpora in low resource languages such as Hindi.
Approach: They propose a Hindi Legal Documents Corpus (HLDC) that contains 900K legal documents in Hindi.
Outcome: The proposed model is based on a corpus of more than 900K legal documents in Hindi.
A Survey on Cross-Lingual Summarization (2022.tacl-1)

Copied to clipboard

Challenge: Cross-lingual summarization is a task of generating a summary in one language for a given document in a different language.
Approach: They present a systematic review of the literature on cross-lingual summarization . they summarize previous efforts and compare them with each other .
Outcome: The proposed approach is compared with previous approaches and summarizes them to provide a deeper analysis.
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
MASALA: Modelling and Analysing the Semantics of Adpositions in Linguistic Annotation of Hindi (2022.lrec-1)

Copied to clipboard

Challenge: Existing work on SNACS annotation for a variety of typologically diverse languages focuses on semantic role labelling and upstream applications in related languages.
Approach: They propose to use the multilingual SNACS annotation scheme to attempt automatic labelling of SNAC supersenses in Hindi.
Outcome: The proposed method is competitive with previous work on English and Gujarati.
CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1,500+ Language Pairs (2023.acl-long)

Copied to clipboard

Challenge: a large-scale cross-lingual summarization dataset is available for free . a cross-linguistic summarizing model can be trained in any target language .
Approach: They propose a multistage data sampling algorithm to train a cross-lingual summarization model capable of summarizing an article in any target language.
Outcome: The proposed model outperforms baseline models on ROUGE and LaSE.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations