Challenge: Visual storytelling aims to automatically generate a coherent story based on a given image sequence.
Approach: They propose a framework that represents the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge.
Outcome: The proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations.

Similar Papers

Plot and Rework: Modeling Storylines for Visual Storytelling (2021.findings-acl)

Copied to clipboard

Challenge: Automated visual storytelling models do not make extensive use of external knowledge and iterative generation when attempting to create stories.
Approach: They propose a framework that uses an image sequence as a story graph to create a coherent story.
Outcome: The proposed framework produces stories superior in diversity, coherence, and humanness . it uses plotting and reworking to improve the model's performance, the authors say .
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for visual storytelling ignore latent topic information.
Approach: They propose a topic-aware reinforcement network for VIsual StoryTelling that takes topic information into account to generate a coherent story.
Outcome: The proposed method outperforms most of the competing models across multiple evaluation metrics.
Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for visual storytelling construct text description independently for each image and roughly concatenate them as a story, which leads to the problem of generating semantically incoherent content.
Approach: They propose a topic description task to detect the global semantic context of an image stream and a story is then constructed with the guidance of the topic description.
Outcome: The proposed framework can generate stories with higher quality compared to state-of-the-art methods on a VIST dataset.
RoViST: Learning Robust Metrics for Visual Storytelling (2022.findings-naacl)

Copied to clipboard

Challenge: Visual storytelling is the task of generating a story paragraph that describes a given image sequence.
Approach: They propose 3 evaluation metrics sets that analyze which aspects we would look for in a good story . they compare their correlation with human judgement scores on a sample of machine stories .
Outcome: The proposed evaluation metrics outperform other metrics on human correlation on a sample of machine stories from state-of-the-art models.
Integrating Visuospatial, Linguistic, and Commonsense Structure into Story Visualization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work on text-to-image synthesis does not explore the use of linguistic structure of the input text.
Approach: They propose to use constituency parse trees to encode structured input and a Transformer-based recurrent architecture to augment commonsense information to generate visual story.
Outcome: The proposed model improves visual quality and consistency of images from a target domain without fine-tuning.
Visual Storytelling with Question-Answer Plans (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models focus on enhancing the representation of image sequences, but the stories are repetitive, illogical, and lacking in detail.
Approach: They propose a framework which integrates visual representations with pretrained language models and planning.
Outcome: The proposed framework combines visual representations with pretrained language models and planning.
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for storytelling lack coherence and consistency, compromising the overall storytelling experience.
Approach: They propose a novel approach that improves the coherence and consistency of automatically generated stories by managing plot nodes and enabling dynamic interactions between different parts of the story.
Outcome: The proposed approach outperforms existing methods in 84.33% of the trials.
V-ALPHASOCIAL: Benchmark and Self-Reflective Chain-of-Thought Generation for Visual Social Commonsense Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Social commonsense reasoning is a multimodal task that requires both textual and visual cues.
Approach: They propose a method that integrates visual cues into social commonsense reasoning tasks.
Outcome: The proposed method improves social commonsense reasoning on a multimodal foundation model.
VISIAR: Empower MLLM for Visual Story Ideation (2025.findings-acl)

Copied to clipboard

Challenge: Existing literature on visual storytelling has not explored the ideation process fully.
Approach: They propose a visual story ideation task that automates the selection and arrangement of visual assets into coherent sequences that convey expressive storylines.
Outcome: The proposed framework surpasses baseline by 33.5% and 18.5%, respectively, on three metrics.
Visual Writing Prompts: Character-Grounded Story Generation with Curated Image Sequences (2023.tacl-1)

Copied to clipboard

Challenge: Existing work on image-based story generation lacks coherent plots for story generation.
Approach: They propose to use image sequences to generate stories from a dataset that has more coherent plots.
Outcome: The proposed model produces more coherent, visually grounded and diverse stories than existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations