| Challenge: | Existing datasets contain only hundreds of samples, resulting in heavy reliance on hand-crafted features or manually annotated data. |
| Approach: | They propose a new domain-specific dataset for multi-document summarization that is 100 times larger than commonly used datasets. |
| Outcome: | The proposed dataset is 100 times larger than commonly used datasets and in another domain than news. |
Similar Papers
OASum: Large-Scale Open Domain Aspect-based Summarization (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing generic summarization methods generate only one summary for all different requests which is not optimal for diverse demands. |
| Approach: | They use crowd-sourced knowledge on Wikipedia to create a large-scale open-domain aspect-based summarization dataset with 1 million different aspects on 2 million Wikipedia pages. |
| Outcome: | The proposed model can generate diverse aspect-based summarizations on Wikipedia with zero/few-shot and fine-tuning on seven downstream datasets. |
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases. |
| Approach: | They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty. |
| Outcome: | The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty. |
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing statistical phrasal or hierarchical machine translation systems relies on a large set of translation rules which results in engineering challenges. |
| Approach: | They propose to use factorized grammar from the field of linguistics as more general translation rules from XTAG English Grammar to generate a manually crafted summarization dataset. |
| Outcome: | The proposed method outperforms existing methods on low-resource language translation tasks with less training data. |
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays (2024.findings-acl)
Copied to clipboard
| Challenge: | Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies. |
| Approach: | They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries. |
| Outcome: | The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries. |
Tell Me Again! a Large-Scale Dataset of Multiple Summaries for the Same Story (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to represent narratives on short-form texts are limited as narrative semantics are an open class. |
| Approach: | They propose to use Wikipedia summaries as a proxy for entire stories or for analysis of the summary itself. |
| Outcome: | The proposed dataset contains 96,831 individual summaries across 29,505 stories. |
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)
Copied to clipboard
| Challenge: | Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems . |
| Approach: | They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language. |
| Outcome: | The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature. |
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks. |
| Approach: | They present MLSUM, the first large-scale MultiLingual SUMmarization dataset. |
| Outcome: | The proposed dataset contains 1.5M+ article/summary pairs in five different languages. |
WikiAsp: A Dataset for Multi-domain Aspect-based Summarization (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing aspects-based summarization models are domain-specific due to large differences in the type of aspects for different domains. |
| Approach: | They propose a large-scale dataset for multi-domain aspect-based summarization using Wikipedia articles from 20 different domains. |
| Outcome: | The proposed model is based on Wikipedia articles from 20 different domains and uses the section titles and boundaries of each article as a proxy for aspect annotation. |
A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal (2020.acl-main)
Copied to clipboard
| Challenge: | Multidocument summarization (MDS) aims to compress large document collections into short summaries. |
| Approach: | They propose a large-scale multidocument summarization dataset that is large both in total number of document clusters and in the size of individual clusters. |
| Outcome: | The proposed dataset is large both in the total number of document clusters and in the size of individual clusters. |
Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles (2020.emnlp-main)
Copied to clipboard
| Challenge: | Multi-XScience is a dataset construction protocol that favours abstractive modeling approaches. |
| Approach: | They propose a large-scale multi-document summarization dataset that is based on articles and lexical databases and WordNet synonymy information to generate related-work sections of a paper. |
| Outcome: | The proposed method is based on lexical databases and WordNet synonymy information to write related work sections of a paper based upon their abstract and the articles they reference. |