Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Visual storytelling is a task of generating a story for a sequence of several temporally-ordered images or video frames. |
| Approach: | They propose a method that measures story quality in terms of human likeness regarding three key aspects highlighted in previous work: visual grounding, coherence, and repetitiveness. |
| Outcome: | The proposed method improves on the foundation model LLaVA but only slightly compared to TAPM, a 50-times smaller visual storytelling model. |
Similar Papers
GROOViST: A Metric for Grounding Objects in Visual Storytelling (2023.emnlp-main)
Copied to clipboard
| Challenge: | Several evaluation metrics for visual storytelling do not consider images at all . authors propose a novel evaluation tool that accounts for cross-modal dependencies and temporal misalignments . |
| Approach: | They propose a visual storytelling evaluation tool that evaluates visual grounding . they use cross-modal dependencies, temporal misalignments and human intuitions . |
| Outcome: | The proposed evaluation tool accounts for cross-modal dependencies, temporal misalignments and human intuitions on visual grounding. |
RoViST: Learning Robust Metrics for Visual Storytelling (2022.findings-naacl)
Copied to clipboard
| Challenge: | Visual storytelling is the task of generating a story paragraph that describes a given image sequence. |
| Approach: | They propose 3 evaluation metrics sets that analyze which aspects we would look for in a good story . they compare their correlation with human judgement scores on a sample of machine stories . |
| Outcome: | The proposed evaluation metrics outperform other metrics on human correlation on a sample of machine stories from state-of-the-art models. |
Evaluating Visual Narrative Coherence in Story Visualization via Diversified Storylines (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics and datasets often neglect visual continuity and narrative diversity. |
| Approach: | They propose a visual context-aware metric for story visualization that uses large vision-language models to jointly assess caption fidelity and inter-image consistency. |
| Outcome: | The proposed framework achieves a Spearman’s correlation comparable to human agreement on two benchmarks and blends diverse and controlled narrative elements at adjustable ratios, producing challenging evaluation sets. |
Improving Generation and Evaluation of Visual Stories via Semantic Consistency (2021.naacl-main)
Copied to clipboard
| Challenge: | Story visualization is an underexplored task that requires a generative model to generate images . prior work has focused on image generation but there is room for improvement . |
| Approach: | They propose to add a dual learning framework to reinforce semantic alignment between story and generated images and a copy-transform mechanism to model sequentially-consistent story visualization. |
| Outcome: | The proposed models outperform text-to-image synthesis models on the story visualization task . the proposed models also improve visual quality, coherence and relevance . |
Visual Coherence Loss for Coherent and Visually Grounded Story Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing visual storytelling models fail to generate correct referring expressions for characters, causing 60% of the generated stories to be lacking local coherence. |
| Approach: | They propose a loss function inspired by a linguistic theory of coherence for self-supervised learning for image sequence representations and a feature matching metric to check whether the models generate referring expressions correctly for characters in input image sequences. |
| Outcome: | The proposed features and loss function are effective for generating more coherent and visually grounded stories. |
Plot and Rework: Modeling Storylines for Visual Storytelling (2021.findings-acl)
Copied to clipboard
| Challenge: | Automated visual storytelling models do not make extensive use of external knowledge and iterative generation when attempting to create stories. |
| Approach: | They propose a framework that uses an image sequence as a story graph to create a coherent story. |
| Outcome: | The proposed framework produces stories superior in diversity, coherence, and humanness . it uses plotting and reworking to improve the model's performance, the authors say . |
Tales of Morality: Comparing Human- and LLM-Generated Moral Stories from Visual Cues (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study has found that stories are central to how humans communicate moral values . |
| Approach: | They compare human- and LLM-generated moral narratives based on images annotated by humans for moral content . authors propose a framework for evaluating moral storytelling in vision-language models . |
| Outcome: | The proposed model compared human- and LLM-generated narratives on images . human stories reflect a balanced distribution of moral foundations and coherent narrative arcs, but LLMs emphasize Care foundation and lack emotional resolution. |
Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies on automatic story generation (ASG) rely on human criteria, but there is little research on how well they correlate with human criteria. |
| Approach: | They propose to use human criteria to evaluate automatic story generation (ASG) their paper proposes to use HANNA to quantitatively evaluate correlations between 72 automatic metrics and human criteria. |
| Outcome: | The proposed model compared human criteria with automatic criteria and found that they were significantly better than human criteria. |
Generating Visual Stories with Grounded and Coreferent Characters (2026.tacl-1)
Copied to clipboard
| Challenge: | Experimental results show that our model generates visual stories with consistent and coreferent character mentions compared to baselines and state-of-the-art systems. |
| Approach: | They propose a character-centric approach to visual story generation that uses visual and textual character coreference chains to enrich the VIST benchmark. |
| Outcome: | The proposed model generates visual stories with consistent and coreferent character mentions compared to baselines and state-of-the-art systems. |
STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story Generation (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets lack rich enough contexts to guide models and evaluations are unreliable for long-form creative text. |
| Approach: | They propose a dataset and evaluation platform built from STORIUM . their dataset contains 6K lengthy stories with fine-grained natural language annotations . |
| Outcome: | The proposed model can be used to generate 6K long stories with fine-grained natural language annotations and a user-generated dataset. |