GameWikiSum: a Novel Large Multi-Document Summarization Dataset (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets contain only hundreds of samples, resulting in heavy reliance on hand-crafted features or manually annotated data.
Approach: They propose a new domain-specific dataset for multi-document summarization that is 100 times larger than commonly used datasets.
Outcome: The proposed dataset is 100 times larger than commonly used datasets and in another domain than news.

Similar Papers

OASum: Large-Scale Open Domain Aspect-based Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing generic summarization methods generate only one summary for all different requests which is not optimal for diverse demands.
Approach: They use crowd-sourced knowledge on Wikipedia to create a large-scale open-domain aspect-based summarization dataset with 1 million different aspects on 2 million Wikipedia pages.
Outcome: The proposed model can generate diverse aspect-based summarizations on Wikipedia with zero/few-shot and fine-tuning on seven downstream datasets.
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases.
Approach: They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty.
Outcome: The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty.
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications (2024.naacl-long)

Copied to clipboard

Challenge: Existing statistical phrasal or hierarchical machine translation systems relies on a large set of translation rules which results in engineering challenges.
Approach: They propose to use factorized grammar from the field of linguistics as more general translation rules from XTAG English Grammar to generate a manually crafted summarization dataset.
Outcome: The proposed method outperforms existing methods on low-resource language translation tasks with less training data.
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays (2024.findings-acl)

Copied to clipboard

Challenge: Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies.
Approach: They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries.
Outcome: The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries.
Tell Me Again! a Large-Scale Dataset of Multiple Summaries for the Same Story (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to represent narratives on short-form texts are limited as narrative semantics are an open class.
Approach: They propose to use Wikipedia summaries as a proxy for entire stories or for analysis of the summary itself.
Outcome: The proposed dataset contains 96,831 individual summaries across 29,505 stories.
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)

Copied to clipboard

Challenge: Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems .
Approach: They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language.
Outcome: The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature.
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)

Copied to clipboard

Challenge: Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks.
Approach: They present MLSUM, the first large-scale MultiLingual SUMmarization dataset.
Outcome: The proposed dataset contains 1.5M+ article/summary pairs in five different languages.
WikiAsp: A Dataset for Multi-domain Aspect-based Summarization (2021.tacl-1)

Copied to clipboard

Challenge: Existing aspects-based summarization models are domain-specific due to large differences in the type of aspects for different domains.
Approach: They propose a large-scale dataset for multi-domain aspect-based summarization using Wikipedia articles from 20 different domains.
Outcome: The proposed model is based on Wikipedia articles from 20 different domains and uses the section titles and boundaries of each article as a proxy for aspect annotation.
A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal (2020.acl-main)

Copied to clipboard

Challenge: Multidocument summarization (MDS) aims to compress large document collections into short summaries.
Approach: They propose a large-scale multidocument summarization dataset that is large both in total number of document clusters and in the size of individual clusters.
Outcome: The proposed dataset is large both in the total number of document clusters and in the size of individual clusters.
Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles (2020.emnlp-main)

Copied to clipboard

Challenge: Multi-XScience is a dataset construction protocol that favours abstractive modeling approaches.
Approach: They propose a large-scale multi-document summarization dataset that is based on articles and lexical databases and WordNet synonymy information to generate related-work sections of a paper.
Outcome: The proposed method is based on lexical databases and WordNet synonymy information to write related work sections of a paper based upon their abstract and the articles they reference.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations