SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling (2024.eacl-long)
Copied to clipboard
| Challenge: | Visual storytelling aims to automatically generate a coherent story based on a given image sequence. |
| Approach: | They propose a framework that represents the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge. |
| Outcome: | The proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations. |
Similar Papers
Plot and Rework: Modeling Storylines for Visual Storytelling (2021.findings-acl)
Copied to clipboard
| Challenge: | Automated visual storytelling models do not make extensive use of external knowledge and iterative generation when attempting to create stories. |
| Approach: | They propose a framework that uses an image sequence as a story graph to create a coherent story. |
| Outcome: | The proposed framework produces stories superior in diversity, coherence, and humanness . it uses plotting and reworking to improve the model's performance, the authors say . |
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for visual storytelling ignore latent topic information. |
| Approach: | They propose a topic-aware reinforcement network for VIsual StoryTelling that takes topic information into account to generate a coherent story. |
| Outcome: | The proposed method outperforms most of the competing models across multiple evaluation metrics. |
Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for visual storytelling construct text description independently for each image and roughly concatenate them as a story, which leads to the problem of generating semantically incoherent content. |
| Approach: | They propose a topic description task to detect the global semantic context of an image stream and a story is then constructed with the guidance of the topic description. |
| Outcome: | The proposed framework can generate stories with higher quality compared to state-of-the-art methods on a VIST dataset. |
RoViST: Learning Robust Metrics for Visual Storytelling (2022.findings-naacl)
Copied to clipboard
| Challenge: | Visual storytelling is the task of generating a story paragraph that describes a given image sequence. |
| Approach: | They propose 3 evaluation metrics sets that analyze which aspects we would look for in a good story . they compare their correlation with human judgement scores on a sample of machine stories . |
| Outcome: | The proposed evaluation metrics outperform other metrics on human correlation on a sample of machine stories from state-of-the-art models. |
Integrating Visuospatial, Linguistic, and Commonsense Structure into Story Visualization (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on text-to-image synthesis does not explore the use of linguistic structure of the input text. |
| Approach: | They propose to use constituency parse trees to encode structured input and a Transformer-based recurrent architecture to augment commonsense information to generate visual story. |
| Outcome: | The proposed model improves visual quality and consistency of images from a target domain without fine-tuning. |
Visual Storytelling with Question-Answer Plans (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models focus on enhancing the representation of image sequences, but the stories are repetitive, illogical, and lacking in detail. |
| Approach: | They propose a framework which integrates visual representations with pretrained language models and planning. |
| Outcome: | The proposed framework combines visual representations with pretrained language models and planning. |
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation (2025.findings-acl)
Copied to clipboard
Jiaming Li, Yukun Chen, Ziqiang Liu, Minghuan Tan, Lei Zhang, Yunshui Li, Run Luo, Longze Chen, Jing Luo, Ahmadreza Argha, Hamid Alinejad-Rokny, Wei Zhou, Min Yang
| Challenge: | Existing methods for storytelling lack coherence and consistency, compromising the overall storytelling experience. |
| Approach: | They propose a novel approach that improves the coherence and consistency of automatically generated stories by managing plot nodes and enabling dynamic interactions between different parts of the story. |
| Outcome: | The proposed approach outperforms existing methods in 84.33% of the trials. |
V-ALPHASOCIAL: Benchmark and Self-Reflective Chain-of-Thought Generation for Visual Social Commonsense Reasoning (2025.findings-acl)
Copied to clipboard
Zongyu Lin, Zhikun Xu, Xiaohan Song, Yixin Wan, Xingcheng Yao, Tsung-Han Lin, Selina Song, Pranav Subbaraman, Ben Zhou, Kai-Wei Chang, Yizhou Sun
| Challenge: | Social commonsense reasoning is a multimodal task that requires both textual and visual cues. |
| Approach: | They propose a method that integrates visual cues into social commonsense reasoning tasks. |
| Outcome: | The proposed method improves social commonsense reasoning on a multimodal foundation model. |
VISIAR: Empower MLLM for Visual Story Ideation (2025.findings-acl)
Copied to clipboard
Zhaoyang Xia, Somdeb Sarkhel, Mehrab Tanjim, Stefano Petrangeli, Ishita Dasgupta, Yuxiao Chen, Jinxuan Xu, Di Liu, Saayan Mitra, Dimitris N. Metaxas
| Challenge: | Existing literature on visual storytelling has not explored the ideation process fully. |
| Approach: | They propose a visual story ideation task that automates the selection and arrangement of visual assets into coherent sequences that convey expressive storylines. |
| Outcome: | The proposed framework surpasses baseline by 33.5% and 18.5%, respectively, on three metrics. |
Visual Writing Prompts: Character-Grounded Story Generation with Curated Image Sequences (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing work on image-based story generation lacks coherent plots for story generation. |
| Approach: | They propose to use image sequences to generate stories from a dataset that has more coherent plots. |
| Outcome: | The proposed model produces more coherent, visually grounded and diverse stories than existing models. |