A Repository of Corpora for Summarization (L18-1)

Copied to clipboard

Challenge: Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task.
Approach: They propose a repository containing corpora available to train and evaluate automatic summarization systems.
Outcome: The proposed system is based on a repository of corpora available for summarization tasks.

Similar Papers

Summarization Beyond News: The Automatically Acquired Fandom Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Abstractive summarization methods require large corpora to train neural architectures.
Approach: They propose a novel automatic corpus construction approach that automatically constructs large open-licensed summarization corpora from existing large text collections and an evaluation process with human annotators.
Outcome: The proposed approach can be used to train abstractive summarization models on large corpora and through a manual evaluation with human annotators.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
Beyond Generic Summarization: A Multi-faceted Hierarchical Summarization Corpus of Large Heterogeneous Data (L18-1)

Copied to clipboard

Challenge: Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader.
Approach: They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically.
Outcome: The proposed method can be used to develop and evaluate hierarchical summarization systems.
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)

Copied to clipboard

Challenge: a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments .
Approach: They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization .
Outcome: The proposed evaluation metrics are inconsistent with existing evaluation protocols.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
A Modular Tool for Automatic Summarization (P19-3)

Copied to clipboard

Challenge: Abstractive automatic summarization methods are supervized, but they require large corpora to perform tasks.
Approach: They propose to use a modular tool for automatic summarization that is as simple as possible for end-users.
Outcome: The proposed tool is open source and written in Java . it could be used as a baseline for future work and evaluate methods on different corpora.
Align then Summarize: Automatic Alignment Methods for Summarization Corpus Creation (2020.lrec-1)

Copied to clipboard

Challenge: Summarizing text is not a straightforward task.
Approach: They propose to use automated transcriptions to generate reports from automatic transcriptions as a dataset for neural summarization.
Outcome: The proposed model improves on publicmeetings corpus on a dataset of aligned public meetings.
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications.
Approach: They analyze available dialogue corpora and propose possible approaches to create new ones.
Outcome: The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed .
Relational Summarization for Corpus Analysis (N18-1)

Copied to clipboard

Challenge: Existing methods for summarizing textual content are often ignored . relationshipal questions are ubiquitous and varied.
Approach: They propose a method which generates a natural language summary of the relationship between two lexical items in a corpus without reference to a knowledge base.
Outcome: The proposed method generates a natural language summary of the relationship between two lexical items in a corpus without reference to a knowledge base.
BillSum: A Corpus for Automatic Summarization of US Legislation (D19-54)

Copied to clipboard

Challenge: In the US Congress, over 10,000 bills are introduced each year, with state legislatures introducing tens of thousands of bills.
Approach: They introduce the first dataset for summarizing US Congressional and California state bills . they demonstrate that models built on Congressional bills can be used to summarize California billa .
Outcome: The proposed summarization methods can be applied to states without human-written summaries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations