Challenge: Abstractive approaches to extract salient sentences from long documents are not effective due to their size.
Approach: They propose a simple yet effective approach that exploits the discourse information of a document to select salient sections instead of sentences for summary generation.
Outcome: The proposed approach avoids full-text understanding and retains salient information given the length limit.

Similar Papers

Extractive Summarization of Long Documents by Combining Global and Local Context (D19-1)

Copied to clipboard

Challenge: Existing methods for extractive and abstractive summarization are far from human performance.
Approach: They propose a neural single-document extractive summarization model for long documents that incorporates both the global context of the whole document and the local context.
Outcome: The proposed model outperforms previous models on ROUGE-1, ROUGEE-2 and METEOR scores on two datasets of scientific papers.
TSTR: Too Short to Represent, Summarize with Details! Intro-Guided Extended Summary Generation (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for extractive and abstract summarization are limited to short abstracts . however, extended summaries provide detailed information beyond coarse information .
Approach: They propose an extractive summarization tool that utilizes the introductory information of documents as pointers to their salient information.
Outcome: The proposed extractive summarization improves on existing datasets with human-written summaries . the proposed summarizing improves in terms of cohesion and completeness compared to baselines and state-of-the-art .
A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents (N18-2)

Copied to clipboard

Challenge: Existing abstractive summarization models focus on summarizing sentences and short documents.
Approach: They propose a hierarchical encoder that models the discourse structure of a document, and an attentive discourse-aware decoder to generate the summary.
Outcome: The proposed model significantly outperforms state-of-the-art models on two large-scale datasets of scientific papers.
Leveraging Information Bottleneck for Scientific Document Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to extract salient sentences from document are unsupervised and rely on graph-based methods for sentence ranking.
Approach: They propose an unsupervised extractive approach to document level summarization based on the Information Bottleneck principle.
Outcome: The proposed framework can be extended to a multi-view framework by different signals.
Discourse-Aware Unsupervised Summarization for Long Scientific Documents (2021.eacl-main)

Copied to clipboard

Challenge: Existing extractive models for short news summarization are weak, despite recent advances in abstractive summarizing.
Approach: They propose an unsupervised graph-based ranking model that uses a hierarchical graph representation to determine sentence importance.
Outcome: The proposed model outperforms strong unsupervised baselines by wide margins in automatic metrics and human evaluation.
A Summarization System for Scientific Documents (D19-3)

Copied to clipboard

Challenge: a qualitative user study identified the most valuable scenarios for scientific content consumption.
Approach: They propose a system that retrieves and summarizes scientific documents for a given information need.
Outcome: The proposed system ingested 270,000 scientific papers and validated with human experts.
On Extractive and Abstractive Neural Document Summarization with Transformer Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: We present a method to produce abstractive summaries of documents that exceed several thousand words . we compare transformer based methods to extractive methods, but extractive models score higher .
Approach: They propose a method to generate abstractive summaries of documents that exceed several thousand words via neural abstractive summary.
Outcome: The proposed method produces abstractive summaries of documents that exceed several thousand words . it is compared with baseline methods, state-of-the-art models and variants of the proposed method .
SumSurvey: An Abstractive Dataset of Scientific Survey Papers for Long Document Summarization (2024.findings-acl)

Copied to clipboard

Challenge: a growing need for long document summarization datasets with 16k input is causing problems.
Approach: They propose to use a dataset to analyze salient information in long document summarizations.
Outcome: The proposed dataset outperforms existing models and LLMs in the distribution form of salient information and the distribution of salinal information is an indicator of quality.
Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries (2022.coling-1)

Copied to clipboard

Challenge: Existing models focus on local word prediction, and cannot make high level plans on what to generate.
Approach: They propose a pipelined system that summarises, outlines and elaborates on each bullet point to generate the corresponding segment.
Outcome: The proposed system produces long texts with significantly better quality and faster convergence speed.
SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to handle long text are limited due to time and memory complexity and limited input lengths.
Approach: They propose a multi-stage split-then-summarize framework for long input summarization . their framework can process input text of arbitrary length by adjusting the number of stages .
Outcome: The proposed framework outperforms existing methods on three long meeting summarization datasets and on a long document summarizing dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations