Challenge: Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies.
Approach: They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries.
Outcome: The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries.

Similar Papers

SummScreen: A Dataset for Abstractive Screenplay Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing summarization datasets are constructed from various domains, such as news, and we characterize them using two entity-centric metrics.
Approach: They propose to use a summarization dataset to evaluate TV series transcripts and recaps . they propose to employ two entity-centric metrics to evaluate the dataset .
Outcome: The proposed model outperforms the existing model and its oracle counterparts in character overlap and accuracy.
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases.
Approach: They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty.
Outcome: The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty.
NarraSum: A Large-Scale Dataset for Abstractive Narrative Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on summarizing news documents or structured documents.
Approach: They propose to use a large-scale narrative summarization dataset to encourage research . they find there is a performance gap between humans and the models on NarraSum .
Outcome: The proposed dataset shows that humans and state-of-the-art models perform poorly when summarizing a narrative . it contains 122K narratives collected from synopses of movies and TV episodes with diverse genres .
AligNarr: Aligning Narratives on Movies (2021.acl-short)

Copied to clipboard

Challenge: Experimental results show the viability of an unsupervised approach to align movie scripts with plot summaries.
Approach: They propose an unsupervised method to align movie scripts with plot summaries using a global optimization model.
Outcome: The proposed method outperforms a baseline alignment model on ten movies with 76% F1 score.
SumTitles: a Summarization Dataset with Low Extractiveness (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for extractive summarization of dialogue data are limited by the grammar and structure of the utterances used.
Approach: They propose a low-extractive corpus of movie dialogues for abstractive text summarization . they use an alignment algorithm to construct the corpus and a baseline evaluation .
Outcome: The proposed method is low-extractive and shows high performance in dialogue datasets.
DiscoGraMS: Enhancing Movie Screen-Play Summarization using Movie Character-Aware Discourse Graph (2025.naacl-short)

Copied to clipboard

Challenge: Recent attempts at screenplay summarization focus on fine-tuning transformer-based pre-trained models, but these models often fall short in capturing long-term dependencies and latent relationships.
Approach: They propose a novel resource that represents movie scripts as a movie character-aware discourse graph (CaD Graph) this resource aims to preserve all salient information, offering a more comprehensive and faithful representation of the screenplay’s content.
Outcome: The proposed model preserves all salient information, offering a more comprehensive and faithful representation of the screenplay’s content.
Select and Summarize: Scene Saliency for Movie Script Summarization (2024.findings-naacl)

Copied to clipboard

Challenge: Existing models for summarizing long-form narrative texts are computationally and memory limited.
Approach: They propose a scene saliency dataset that consists of human-annotated salient scenes for 100 movies.
Outcome: The proposed model outperforms state-of-the-art models and reflects the information content of a movie more accurately than a model that takes the whole movie script as input.
Character Coreference Resolution in Movie Screenplays (2023.findings-acl)

Copied to clipboard

Challenge: Movie screenplays have a distinct narrative structure.
Approach: They develop a method to extract structural information and character coreference clusters from movie screenplays by leveraging a movie parser and a character coreferser.
Outcome: The proposed methods scale to long movie screenplays without dramatically increasing their memory footprints.
BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization (P19-1)

Copied to clipboard

Challenge: Existing text summarization datasets are compiled from news articles, where summary-worthy content often appears in the beginning of input articles.
Approach: They present a novel dataset, BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries.
Outcome: The proposed dataset is compared with existing summarization datasets and demonstrates that salient content is evenly distributed in the input.
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)

Copied to clipboard

Challenge: Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks.
Approach: They present MLSUM, the first large-scale MultiLingual SUMmarization dataset.
Outcome: The proposed dataset contains 1.5M+ article/summary pairs in five different languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations