Echoes from Alexandria: A Large Resource for Multilingual Book Summarization (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent research in text summarization has focused on news stories, where texts are typically short and have strong layout features. |
| Approach: | They propose a resource for multilingual book summarization that uses a new extractive-then-abstractive baseline to compare the results. |
| Outcome: | The proposed resource is the largest and first to be multilingual, featuring 5 languages and 25 language pairs. |
Similar Papers
ACLSum: A New Dataset for Aspect-based Summarization of Scientific Publications (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing statistical phrasal or hierarchical machine translation systems relies on a large set of translation rules which results in engineering challenges. |
| Approach: | They propose to use factorized grammar from the field of linguistics as more general translation rules from XTAG English Grammar to generate a manually crafted summarization dataset. |
| Outcome: | The proposed method outperforms existing methods on low-resource language translation tasks with less training data. |
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases. |
| Approach: | They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty. |
| Outcome: | The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty. |
WikiAsp: A Dataset for Multi-domain Aspect-based Summarization (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing aspects-based summarization models are domain-specific due to large differences in the type of aspects for different domains. |
| Approach: | They propose a large-scale dataset for multi-domain aspect-based summarization using Wikipedia articles from 20 different domains. |
| Outcome: | The proposed model is based on Wikipedia articles from 20 different domains and uses the section titles and boundaries of each article as a proxy for aspect annotation. |
Exploring Content Selection in Summarization of Novel Chapters (2020.acl-main)
Copied to clipboard
| Challenge: | We focus on extractive summarization, which requires the creation of a gold-standard set of extractive summary summaries. |
| Approach: | They propose a new metric for aligning summary sentences with chapter sentences to create gold extracts. |
| Outcome: | The proposed method improves on previous methods and automatic metrics and a crowd-sourced pyramid analysis. |
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages (2021.findings-acl)
Copied to clipboard
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, Rifat Shahriyar
| Challenge: | XL-Sum dataset covers 44 languages ranging from low to high-resource . Xl-SUM is highly abstractive, concise, and of high quality . |
| Approach: | They present a dataset comprising 1 million professionally annotated article-summary pairs from BBC . they fine-tune a pretrained multilingual model with XL-Sum and experiment on multilingual and lowresource tasks. |
| Outcome: | The proposed dataset is highly abstractive, concise, and of high quality . it shows higher scores on 10 languages than similar datasets compared to monolingual ones . |
LR-Sum: Summarization for Less-Resourced Languages (2023.findings-acl)
Copied to clipboard
| Challenge: | LR-Sum contains human-written summaries for 40 languages, many of which are less-resourced. |
| Approach: | They propose to use a permissively-licensed dataset to analyze human-written summaries for 40 languages. |
| Outcome: | The proposed dataset contains human-written summaries for 40 languages . authors describe abstractive and extractive summarization experiments . |
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset (2021.emnlp-main)
Copied to clipboard
| Challenge: | Current research in automatic summarisation is expensive to create, posing a challenge for any language. |
| Approach: | They propose to use a large-scale multilingual summarisation dataset with articles in 92 languages and more than 35 writing scripts to generate a multilingual dataset. |
| Outcome: | The proposed method is the largest, most inclusive, existing dataset and one of the largest and most inclusive datasets for any NLP task. |
Models and Datasets for Cross-Lingual Summarisation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent years have witnessed increased interest in abstractive summarisation thanks to the popularity of neural network models and the availability of datasets containing hundreds of thousands of document-summary pairs. |
| Approach: | They propose to create a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in . target language. |
| Outcome: | The proposed task can be applied to several other languages and covers twelve languages and directions. |
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs (2026.acl-long)
Copied to clipboard
Abdellah EL Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah A. Alsayadi, Nadia Ghezaiel Hammouda, Alshima Mohammed Alkhazimi, Aya Hamod, Al-Yas Yaqoob Al-Ghafri, Wesam El-Sayed, Asila Ismail al Sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Aeej Mohammed Aseri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Emehah, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed S. Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Said Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Al-kathiri, Fadi Zaraket, Mustafa Jarrar, Yahya Mohamed EL Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed
| Challenge: | Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than modern standard Arabic (MSA). |
| Approach: | They propose a large-scale, community-driven, human-translated dataset to bridge this gap . Alexandria covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . |
| Outcome: | The Alexandria dataset covers 13 Arab countries and 11 high-impact domains . it provides unprecedented granularity by associating contributions with city-of-origin metadata . Alexandria is a training resource and a rigorous benchmark for evaluating MT and LLMs based on the Alexandria dataset . |
Beyond Generic Summarization: A Multi-faceted Hierarchical Summarization Corpus of Large Heterogeneous Data (L18-1)
Copied to clipboard
| Challenge: | Automated summarization has focused on ten to twenty documents, typically news articles, but could in theory analyze hundreds of documents from a wide range of sources and provide an overview to the interested reader. |
| Approach: | They propose a method for creating hierarchical summarization corpora from large, heterogeneous document collections by crowdsourcing relevant content and asking trained annotators to order the relevant information hierarchically. |
| Outcome: | The proposed method can be used to develop and evaluate hierarchical summarization systems. |