Challenge: a dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications . identifying large, high-quality resources for summarization has called for creative solutions in the past.
Approach: They present a summarization dataset of 1.3 million articles and summaries written by newsrooms of 38 major news publications.
Outcome: The summarization dataset shows high diversity of summarizing styles . authors train existing methods on the data to evaluate its utility and challenges.

Similar Papers

BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization (P19-1)

Copied to clipboard

Challenge: Existing text summarization datasets are compiled from news articles, where summary-worthy content often appears in the beginning of input articles.
Approach: They present a novel dataset, BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries.
Outcome: The proposed dataset is compared with existing summarization datasets and demonstrates that salient content is evenly distributed in the input.
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset (2021.emnlp-main)

Copied to clipboard

Challenge: Current research in automatic summarisation is expensive to create, posing a challenge for any language.
Approach: They propose to use a large-scale multilingual summarisation dataset with articles in 92 languages and more than 35 writing scripts to generate a multilingual dataset.
Outcome: The proposed method is the largest, most inclusive, existing dataset and one of the largest and most inclusive datasets for any NLP task.
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)

Copied to clipboard

Challenge: Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks.
Approach: They present MLSUM, the first large-scale MultiLingual SUMmarization dataset.
Outcome: The proposed dataset contains 1.5M+ article/summary pairs in five different languages.
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases.
Approach: They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty.
Outcome: The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty.
DaNewsroom: A Large-scale Danish Summarisation Dataset (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets for automatic summarisation are English-oriented . however, only very limited datasets exist in languages other than English .
Approach: They present the first large-scale non-English dataset specifically curated for automatic summarisation.
Outcome: The proposed dataset is the first for the Danish language and is compared with existing datasets.
Multi-News: A Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model (P19-1)

Copied to clipboard

Challenge: Multi-document summarization (MDS) of news articles has been limited to datasets of a couple of hundred examples.
Approach: They propose a model which integrates a traditional extractive summarization model with a standard SDS model and achieves competitive results on MDS datasets.
Outcome: The proposed model achieves competitive results on large-scale datasets.
Exploring Content Selection in Summarization of Novel Chapters (2020.acl-main)

Copied to clipboard

Challenge: We focus on extractive summarization, which requires the creation of a gold-standard set of extractive summary summaries.
Approach: They propose a new metric for aligning summary sentences with chapter sentences to create gold extracts.
Outcome: The proposed method improves on previous methods and automatic metrics and a crowd-sourced pyramid analysis.
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies on multi-document summarization focus on collating information that all sources agree upon, but the task of summarizing diverse information remains underexplored.
Approach: They propose a task of summarizing diverse information encountered in multiple news articles encompassing the same event using a dataset curated by a large language model.
Outcome: The proposed task aims to summarize diverse information in multiple news articles encompassing the same event . the proposed task is difficult due to its limited coverage and verbosity biases .
SumPubMed: Summarization Dataset of PubMed Scientific Articles (2021.acl-srw)

Copied to clipboard

Challenge: Existing summarization models that can extract the top few lines of news articles fail to summarize long documents.
Approach: They constructed a scientific summarization dataset from MEDLINE articles from the PubMed archive to address this problem.
Outcome: The proposed model outperforms existing models on news article summarization datasets and shows that it is more efficient to extract the top few lines.
CCSum: A Large-Scale and High-Quality Dataset for Abstractive News Summarization (2024.naacl-long)

Copied to clipboard

Challenge: Existing datasets for supervised news summarization contain considerable amount of noise and expensive training data.
Approach: They propose a large-scale and high-quality dataset for supervised abstractive news summarization containing 1.3 million training samples.
Outcome: The proposed dataset is more factual and informative than established summarization datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations