BillSum: A Corpus for Automatic Summarization of US Legislation (D19-54)

Copied to clipboard

Challenge: In the US Congress, over 10,000 bills are introduced each year, with state legislatures introducing tens of thousands of bills.
Approach: They introduce the first dataset for summarizing US Congressional and California state bills . they demonstrate that models built on Congressional bills can be used to summarize California billa .
Outcome: The proposed summarization methods can be applied to states without human-written summaries.

Similar Papers

A Repository of Corpora for Summarization (L18-1)

Copied to clipboard

Challenge: Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task.
Approach: They propose a repository containing corpora available to train and evaluate automatic summarization systems.
Outcome: The proposed system is based on a repository of corpora available for summarization tasks.
Summarization Beyond News: The Automatically Acquired Fandom Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Abstractive summarization methods require large corpora to train neural architectures.
Approach: They propose a novel automatic corpus construction approach that automatically constructs large open-licensed summarization corpora from existing large text collections and an evaluation process with human annotators.
Outcome: The proposed approach can be used to train abstractive summarization models on large corpora and through a manual evaluation with human annotators.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions (2025.findings-naacl)

Copied to clipboard

Challenge: CaseSumm is a dataset for long-context summarization in the legal domain . human groundtruth summaries are often not available for legal summarizing .
Approach: They propose a dataset for long-context summarization that includes SCOTUS opinions and their official summaries.
Outcome: The proposed dataset is the largest open legal case summarization dataset . it outperforms larger models on automatic metrics and human evaluation .
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases.
Approach: They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty.
Outcome: The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty.
Analysis of State-Level Legislative Process in Enhanced Linguistic and Nationwide Network Contexts (2024.naacl-long)

Copied to clipboard

Challenge: a new framework for understanding state-level legislative process improves understanding of state legislation and its implications.
Approach: They propose to use generative large language models to decode legislators' behavior and implications of state policies by establishing a shared nationwide network.
Outcome: The framework decodes legislators’ behavior and implications of state policies by establishing a shared nationwide network enriched with diverse contexts, such as information on interest groups influencing public policy and legislators' courage test results, which reflect their political positions.
Auto-hMDS: Automatic Construction of a Large Heterogeneous Multilingual Multi-Document Summarization Corpus (L18-1)

Copied to clipboard

Challenge: Existing datasets for automatic text summarization are small and focused on newswires.
Approach: They propose to automatically generate a large multilingual multi-document summarization corpus using Wikipedia articles as summaries and to automatically search for appropriate source documents.
Outcome: The proposed corpus contains 7,316 topics in English and German with different summary lengths and number of source documents.
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays (2024.findings-acl)

Copied to clipboard

Challenge: Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies.
Approach: They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries.
Outcome: The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries.
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)

Copied to clipboard

Challenge: Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems .
Approach: They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language.
Outcome: The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature.
Beyond Generic Summarization: A Multi-faceted Hierarchical Summarization Corpus of Large Heterogeneous Data (L18-1)

Copied to clipboard

Challenge: Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader.
Approach: They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically.
Outcome: The proposed method can be used to develop and evaluate hierarchical summarization systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations