A Summarization System for Scientific Documents (D19-3)

Copied to clipboard

Challenge: a qualitative user study identified the most valuable scenarios for scientific content consumption.
Approach: They propose a system that retrieves and summarizes scientific documents for a given information need.
Outcome: The proposed system ingested 270,000 scientific papers and validated with human experts.

Similar Papers

Enhancing Scientific Document Summarization with Research Community Perspective and Background Knowledge (2024.lrec-main)

Copied to clipboard

Challenge: Scientific paper summarization is the focus of recent research . prevailing summarizing methods involve selective extraction of content from abstract, introduction, and conclusion segments within the target articles.
Approach: They propose a model that incorporates references and citations to capture the impact of the document on the research community.
Outcome: The proposed model generates extractive and abstractive summaries in parallel and improves their performance when considering the standard metrics.
SumPubMed: Summarization Dataset of PubMed Scientific Articles (2021.acl-srw)

Copied to clipboard

Challenge: Existing summarization models that can extract the top few lines of news articles fail to summarize long documents.
Approach: They constructed a scientific summarization dataset from MEDLINE articles from the PubMed archive to address this problem.
Outcome: The proposed model outperforms existing models on news article summarization datasets and shows that it is more efficient to extract the top few lines.
Bringing Structure into Summaries: a Faceted Summarization Dataset for Long Scientific Documents (2021.acl-short)

Copied to clipboard

Challenge: Faceted summarization provides briefings of a document from different perspectives.
Approach: They propose a faceted summarization benchmark built on Emerald journal articles . they propose faceted models that bring structure into faceted documents .
Outcome: The proposed benchmark is based on Emerald journal articles and covers a diverse range of domains.
Leveraging Information Bottleneck for Scientific Document Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to extract salient sentences from document are unsupervised and rely on graph-based methods for sentence ranking.
Approach: They propose an unsupervised extractive approach to document level summarization based on the Information Bottleneck principle.
Outcome: The proposed framework can be extended to a multi-view framework by different signals.
Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles (2020.emnlp-main)

Copied to clipboard

Challenge: Multi-XScience is a dataset construction protocol that favours abstractive modeling approaches.
Approach: They propose a large-scale multi-document summarization dataset that is based on articles and lexical databases and WordNet synonymy information to generate related-work sections of a paper.
Outcome: The proposed method is based on lexical databases and WordNet synonymy information to write related work sections of a paper based upon their abstract and the articles they reference.
ROUGE-SciQFS: A ROUGE-based Method to Automatically Create Datasets for Scientific Query-Focused Summarization (2025.coling-main)

Copied to clipboard

Challenge: Scientific Query-Focused Summarization (Sci-QFS) has lagged in development due to the lack of data.
Approach: They propose a method to take advantage of existing academic papers to obtain large-scale datasets for this task automatically.
Outcome: The proposed method outperforms existing models on the datasets and shows that it is relatively straightforward for humans.
The State and Fate of Summarization Datasets: A Survey (2025.naacl-long)

Copied to clipboard

Challenge: Summarization is the task of shortening a text while preserving the most important information it contains.
Approach: They propose a novel ontology covering sample properties, collection methods and distribution covering sample characteristics, collection method and distribution.
Outcome: The proposed ontology covers sample properties, collection methods and distribution, and can be used to streamline future research into a more coherent body of work.
TalkSumm: A Dataset and Scalable Annotation Method for Scientific Paper Summarization Based on Conference Talks (P19-1)

Copied to clipboard

Challenge: Currently, no large-scale training data is available for the task of scientific paper summarization.
Approach: They propose a method that automatically generates scientific paper summaries by utilizing videos of scientific conferences.
Outcome: The proposed model performs similar to models trained on a dataset of summaries created manually.
Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for lay summarisation are limited in size and scope, hindering the development of data-driven approaches.
Approach: They propose to use two new datasets for the lay summarisation of biomedical research articles to characterise their lay summaries.
Outcome: The proposed datasets are compared with existing datasets and show they can be leveraged to support different audiences and applications.
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)

Copied to clipboard

Challenge: Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems .
Approach: They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language.
Outcome: The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations