Challenge: Visual storytelling is a task of generating a story for a sequence of several temporally-ordered images or video frames.
Approach: They propose a method that measures story quality in terms of human likeness regarding three key aspects highlighted in previous work: visual grounding, coherence, and repetitiveness.
Outcome: The proposed method improves on the foundation model LLaVA but only slightly compared to TAPM, a 50-times smaller visual storytelling model.

Similar Papers

GROOViST: A Metric for Grounding Objects in Visual Storytelling (2023.emnlp-main)

Copied to clipboard

Challenge: Several evaluation metrics for visual storytelling do not consider images at all . authors propose a novel evaluation tool that accounts for cross-modal dependencies and temporal misalignments .
Approach: They propose a visual storytelling evaluation tool that evaluates visual grounding . they use cross-modal dependencies, temporal misalignments and human intuitions .
Outcome: The proposed evaluation tool accounts for cross-modal dependencies, temporal misalignments and human intuitions on visual grounding.
RoViST: Learning Robust Metrics for Visual Storytelling (2022.findings-naacl)

Copied to clipboard

Challenge: Visual storytelling is the task of generating a story paragraph that describes a given image sequence.
Approach: They propose 3 evaluation metrics sets that analyze which aspects we would look for in a good story . they compare their correlation with human judgement scores on a sample of machine stories .
Outcome: The proposed evaluation metrics outperform other metrics on human correlation on a sample of machine stories from state-of-the-art models.
Evaluating Visual Narrative Coherence in Story Visualization via Diversified Storylines (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation metrics and datasets often neglect visual continuity and narrative diversity.
Approach: They propose a visual context-aware metric for story visualization that uses large vision-language models to jointly assess caption fidelity and inter-image consistency.
Outcome: The proposed framework achieves a Spearman’s correlation comparable to human agreement on two benchmarks and blends diverse and controlled narrative elements at adjustable ratios, producing challenging evaluation sets.
Improving Generation and Evaluation of Visual Stories via Semantic Consistency (2021.naacl-main)

Copied to clipboard

Challenge: Story visualization is an underexplored task that requires a generative model to generate images . prior work has focused on image generation but there is room for improvement .
Approach: They propose to add a dual learning framework to reinforce semantic alignment between story and generated images and a copy-transform mechanism to model sequentially-consistent story visualization.
Outcome: The proposed models outperform text-to-image synthesis models on the story visualization task . the proposed models also improve visual quality, coherence and relevance .
Visual Coherence Loss for Coherent and Visually Grounded Story Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing visual storytelling models fail to generate correct referring expressions for characters, causing 60% of the generated stories to be lacking local coherence.
Approach: They propose a loss function inspired by a linguistic theory of coherence for self-supervised learning for image sequence representations and a feature matching metric to check whether the models generate referring expressions correctly for characters in input image sequences.
Outcome: The proposed features and loss function are effective for generating more coherent and visually grounded stories.
Plot and Rework: Modeling Storylines for Visual Storytelling (2021.findings-acl)

Copied to clipboard

Challenge: Automated visual storytelling models do not make extensive use of external knowledge and iterative generation when attempting to create stories.
Approach: They propose a framework that uses an image sequence as a story graph to create a coherent story.
Outcome: The proposed framework produces stories superior in diversity, coherence, and humanness . it uses plotting and reworking to improve the model's performance, the authors say .
Tales of Morality: Comparing Human- and LLM-Generated Moral Stories from Visual Cues (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study has found that stories are central to how humans communicate moral values .
Approach: They compare human- and LLM-generated moral narratives based on images annotated by humans for moral content . authors propose a framework for evaluating moral storytelling in vision-language models .
Outcome: The proposed model compared human- and LLM-generated narratives on images . human stories reflect a balanced distribution of moral foundations and coherent narrative arcs, but LLMs emphasize Care foundation and lack emotional resolution.
Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story Generation (2022.coling-1)

Copied to clipboard

Challenge: Existing studies on automatic story generation (ASG) rely on human criteria, but there is little research on how well they correlate with human criteria.
Approach: They propose to use human criteria to evaluate automatic story generation (ASG) their paper proposes to use HANNA to quantitatively evaluate correlations between 72 automatic metrics and human criteria.
Outcome: The proposed model compared human criteria with automatic criteria and found that they were significantly better than human criteria.
Generating Visual Stories with Grounded and Coreferent Characters (2026.tacl-1)

Copied to clipboard

Challenge: Experimental results show that our model generates visual stories with consistent and coreferent character mentions compared to baselines and state-of-the-art systems.
Approach: They propose a character-centric approach to visual story generation that uses visual and textual character coreference chains to enrich the VIST benchmark.
Outcome: The proposed model generates visual stories with consistent and coreferent character mentions compared to baselines and state-of-the-art systems.
STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing datasets lack rich enough contexts to guide models and evaluations are unreliable for long-form creative text.
Approach: They propose a dataset and evaluation platform built from STORIUM . their dataset contains 6K lengthy stories with fine-grained natural language annotations .
Outcome: The proposed model can be used to generate 6K long stories with fine-grained natural language annotations and a user-generated dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations