Challenge: Currently, no large-scale training data is available for the task of scientific paper summarization.
Approach: They propose a method that automatically generates scientific paper summaries by utilizing videos of scientific conferences.
Outcome: The proposed model performs similar to models trained on a dataset of summaries created manually.

Similar Papers

What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations (2025.acl-long)

Copied to clipboard

Challenge: VISTA dataset contains 18,599 recorded AI conference presentations . large multimodal models exhibit reduced performance in scientific contexts, study shows .
Approach: They propose a dataset specifically designed for video-to-text summarization in scientific domains.
Outcome: This paper compares the performance of large models with human models and shows that they improve on human models.
A Summarization System for Scientific Documents (D19-3)

Copied to clipboard

Challenge: a qualitative user study identified the most valuable scenarios for scientific content consumption.
Approach: They propose a system that retrieves and summarizes scientific documents for a given information need.
Outcome: The proposed system ingested 270,000 scientific papers and validated with human experts.
ForumSum: A Multi-Speaker Conversation Summarization Dataset (2021.findings-emnlp)

Copied to clipboard

Challenge: Abstractive summarization quality has been improved but there is a lack of data for conversation summarizing applications.
Approach: They propose to build a conversation summarization dataset with human written summaries from internet forums.
Outcome: The proposed dataset can be easily expanded to improve conversation summarization applications.
CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited Supervision (2022.emnlp-main)

Copied to clipboard

Challenge: Scientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers . previous attempts failed to scale up due to heavy human annotation and domain expertise .
Approach: They propose a method to automatically extract TLDR summaries from scientific papers . they propose 'citeSum' with no human annotation, which is 30 times larger than SciTLDR .
Outcome: The proposed approach outperforms most fully-supervised methods on SciTLDR without fine-tuning and achieves state-of-the-art results with only 128 examples.
TLDR: Extreme Summarization of Scientific Documents (2020.findings-emnlp)

Copied to clipboard

Challenge: TLDR generation requires expert background knowledge and understanding of complex domain-specific language.
Approach: They propose a learning strategy that exploits titles as an auxiliary training signal.
Outcome: The proposed method improves upon strong baselines under both automated metrics and human evaluations.
Enhancing Scientific Document Summarization with Research Community Perspective and Background Knowledge (2024.lrec-main)

Copied to clipboard

Challenge: Scientific paper summarization is the focus of recent research . prevailing summarizing methods involve selective extraction of content from abstract, introduction, and conclusion segments within the target articles.
Approach: They propose a model that incorporates references and citations to capture the impact of the document on the research community.
Outcome: The proposed model generates extractive and abstractive summaries in parallel and improves their performance when considering the standard metrics.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing work on abstractive dialogue summarizations has focused on news summarizing but there is no such comprehensive dataset.
Approach: They propose to use a chat-dialogues corpus with abstractive dialogue summaries to generate a short version of text that covers the main points succinctly.
Outcome: The proposed dataset achieves higher ROUGE scores than the model-generated summaries of news, compared with human evaluators' judgement.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
Longform Multimodal Lay Summarization of Scientific Papers: Towards Automatically Generating Science Blogs from Research Articles (2024.lrec-main)

Copied to clipboard

Challenge: Science blogs and lay-speak are critical to communicating scientific information to the general public and policymakers.
Approach: They propose to use presentation transcripts and slides to generate a scientific blog from a research article in layperson's terms.
Outcome: The proposed approach can generate a blog text and select the most relevant figures to explain a research article in layperson’s terms, essentially a science blog.
SumPubMed: Summarization Dataset of PubMed Scientific Articles (2021.acl-srw)

Copied to clipboard

Challenge: Existing summarization models that can extract the top few lines of news articles fail to summarize long documents.
Approach: They constructed a scientific summarization dataset from MEDLINE articles from the PubMed archive to address this problem.
Outcome: The proposed model outperforms existing models on news article summarization datasets and shows that it is more efficient to extract the top few lines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations