Liputan6: A Large-scale Indonesian Dataset for Text Summarization (2020.aacl-main)
Copied to clipboard
| Challenge: | Despite having the fourth largest speaker population in the world, 1 Indonesian is under-represented in NLP. |
| Approach: | They propose to use a large-scale Indonesian summarization dataset to test extractive and abstractive summarizing methods. |
| Outcome: | The proposed methods are compared with multilingual and monolingual BERT-based models. |
Similar Papers
BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization (P19-1)
Copied to clipboard
| Challenge: | Existing text summarization datasets are compiled from news articles, where summary-worthy content often appears in the beginning of input articles. |
| Approach: | They present a novel dataset, BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries. |
| Outcome: | The proposed dataset is compared with existing summarization datasets and demonstrates that salient content is evenly distributed in the input. |
LipKey: A Large-Scale News Dataset for Absent Keyphrases Generation and Abstractive Summarization (2022.coling-1)
Copied to clipboard
| Challenge: | Existing work has addressed each element individually, but this study focuses on LipKey, the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. |
| Approach: | They propose a novel news dataset that consists of highly absent keyphrases . they combine lips keyphrase and TF-IDF to obtain abstractive summaries . |
| Outcome: | The proposed dataset is the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles. |
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages (2021.findings-acl)
Copied to clipboard
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, Rifat Shahriyar
| Challenge: | XL-Sum dataset covers 44 languages ranging from low to high-resource . Xl-SUM is highly abstractive, concise, and of high quality . |
| Approach: | They present a dataset comprising 1 million professionally annotated article-summary pairs from BBC . they fine-tune a pretrained multilingual model with XL-Sum and experiment on multilingual and lowresource tasks. |
| Outcome: | The proposed dataset is highly abstractive, concise, and of high quality . it shows higher scores on 10 languages than similar datasets compared to monolingual ones . |
Large Scale Multi-Lingual Multi-Modal Summarization Dataset (2023.eacl-main)
Copied to clipboard
| Challenge: | a large dataset of document-image pairs and annotated multi-modal summarization data is needed for multi-lingual modeling . encoder-decoder models represent information comprising multiple modalities. |
| Approach: | They propose to use a multi-lingual summarization dataset to analyze multi-modal summarizing using multi-linguistic annotated data. |
| Outcome: | The proposed dataset is the largest multi-lingual multi-modal summarization dataset for 13 languages and consists of cross-lingual summarizing data for 2 languages. |
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities. |
| Approach: | They propose a multimodal article and video summarization dataset that integrates resources from different modalities. |
| Outcome: | The proposed dataset validates the important assistance role of external information for multimodal summarization. |
CATAMARAN: A Cross-lingual Long Text Abstractive Summarization Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on cross-lingual summarization rely on pseudo-cross-lingual datasets . such an approach would lead to the loss of information in the original document and introduce noise into the summary . |
| Approach: | They present a high-quality cross-lingual long text abstractive summarization dataset . it contains 20,000 parallel news articles and corresponding summaries written by humans . |
| Outcome: | The proposed model outperforms monolingual systems in the cross-lingual task. |
A Summarization Dataset of Slovak News Articles (2020.lrec-1)
Copied to clipboard
| Challenge: | a number of studies on document summarization have focused on the English language . however, most of the work on this task is done on English datasets . |
| Approach: | They propose to use a news site's ROUGE metric to adapt it to Slovak texts . they propose to introduce a large-scale news-based summarization dataset . |
| Outcome: | The proposed approach is better suited for Slovak texts than the dominant ROUGE metric. |
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset (2021.emnlp-main)
Copied to clipboard
| Challenge: | Current research in automatic summarisation is expensive to create, posing a challenge for any language. |
| Approach: | They propose to use a large-scale multilingual summarisation dataset with articles in 92 languages and more than 35 writing scripts to generate a multilingual dataset. |
| Outcome: | The proposed method is the largest, most inclusive, existing dataset and one of the largest and most inclusive datasets for any NLP task. |
EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing summarization datasets focus on overly exposed domains and are primarily monolingual with few multilingual datasets. |
| Approach: | They propose a new summarization dataset based on manually curated document summaries from the European Union law platform EUR-Lex. |
| Outcome: | The proposed dataset is based on document summaries of legal acts from the European Union law platform (EUR-Lex). |
Beyond Generic Summarization: A Multi-faceted Hierarchical Summarization Corpus of Large Heterogeneous Data (L18-1)
Copied to clipboard
| Challenge: | Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader. |
| Approach: | They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically. |
| Outcome: | The proposed method can be used to develop and evaluate hierarchical summarization systems. |