Challenge: The SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel, with a lag behind actual broadcast time of at most a few minutes.
Approach: The open-source SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel . it offers a fully automated media ingestion pipeline capable of recording live broadcasts, detection and transcription of spoken content, translation of all text (original or transcribed) into English, recognition and linking of Named Entities, topic detection, clustering and cross-lingual multi-document summarization of related media items and extraction and storage of factual claims in these news items.
Outcome: The SUMMA Platform is a highly scalable distributed architecture for monitoring a large number of media broadcasts in parallel, with a lag behind actual broadcast time of at most a few minutes.

Similar Papers

MassiveSumm: a very large-scale, very multilingual, news summarisation dataset (2021.emnlp-main)

Copied to clipboard

Challenge: Current research in automatic summarisation is expensive to create, posing a challenge for any language.
Approach: They propose to use a large-scale multilingual summarisation dataset with articles in 92 languages and more than 35 writing scripts to generate a multilingual dataset.
Outcome: The proposed method is the largest, most inclusive, existing dataset and one of the largest and most inclusive datasets for any NLP task.
PMIndiaSum: Multilingual and Cross-lingual Headline Summarization for Languages in India (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets for Indian languages are limited in terms of coverage and size.
Approach: They propose a multilingual and massively parallel summarization corpus focused on languages in India that provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
Outcome: The proposed dataset provides a training and testing ground for four language families, 14 languages, and the largest to date with 196 language pairs.
GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Current studies focus on single-language or single-document tasks for news summarization . lack of a benchmark inhibits researchers from adequately studying this invaluable problem.
Approach: They propose a novel task that unifies Multi-lingual, Cross-lingual and Multi-document Summarization into one task.
Outcome: The proposed task encapsulates the real-world requirements all-in-one and is validated by extensive analysis.
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)

Copied to clipboard

Challenge: Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks.
Approach: They present MLSUM, the first large-scale MultiLingual SUMmarization dataset.
Outcome: The proposed dataset contains 1.5M+ article/summary pairs in five different languages.
A Workbench for Rapid Generation of Cross-Lingual Summaries (L18-1)

Copied to clipboard

Challenge: a tool for automating cross-lingual information access is needed in multilingual societies . current state of machine translation is not able to generate publishable articles from English .
Approach: They propose a web-based tool for human editing of cross-lingual summaries . it generates publishable summary in a number of Indian Languages for news articles originally published in english .
Outcome: The proposed tool can generate publishable summaries in multiple languages with minimal human effort and collect detailed logs on the process.
Large Scale Multi-Lingual Multi-Modal Summarization Dataset (2023.eacl-main)

Copied to clipboard

Challenge: a large dataset of document-image pairs and annotated multi-modal summarization data is needed for multi-lingual modeling . encoder-decoder models represent information comprising multiple modalities.
Approach: They propose to use a multi-lingual summarization dataset to analyze multi-modal summarizing using multi-linguistic annotated data.
Outcome: The proposed dataset is the largest multi-lingual multi-modal summarization dataset for 13 languages and consists of cross-lingual summarizing data for 2 languages.
Multilingual Clustering of Streaming News (D18-1)

Copied to clipboard

Challenge: a novel method for clustering news across languages is proposed . a key challenge in handling news streams is that they must be generated on the fly .
Approach: They propose a method for clustering news across languages into monolingual and crosslingual clusters . they use real news datasets in multiple languages to find an ever growing number of cluster labels .
Outcome: The proposed method produces state-of-the-art results on real news datasets in German, English and Spanish.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal Domain (2022.emnlp-main)

Copied to clipboard

Challenge: Existing summarization datasets focus on overly exposed domains and are primarily monolingual with few multilingual datasets.
Approach: They propose a new summarization dataset based on manually curated document summaries from the European Union law platform EUR-Lex.
Outcome: The proposed dataset is based on document summaries of legal acts from the European Union law platform (EUR-Lex).
MARS: Multilingual Aspect-centric Review Summarisation (2024.emnlp-industry)

Copied to clipboard

Challenge: Existing methods for summarizing customer feedback are not able to extract actionable reviews into a specific target language.
Approach: They propose a framework involving extract-then-summarise to summariser customer feedback into a specific language.
Outcome: The proposed framework improves abstractive baselines and efficiency to real-time systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations