Challenge: Story video-text alignment is a core task in computational story understanding, but its progress has been held back by the scarcity of manually annotated video- text correspondences and the heavy concentration on English narrations of Hollywood movies.
Approach: They construct a multilingual video story dataset with 13,166 movie summary videos from 7 languages and manual annotations of fine-grained video-text correspondences.
Outcome: The proposed approach outperforms the SOTA methods on clip accuracy and Sentence IoU scores.

Similar Papers

AligNarr: Aligning Narratives on Movies (2021.acl-short)

Copied to clipboard

Challenge: Experimental results show the viability of an unsupervised approach to align movie scripts with plot summaries.
Approach: They propose an unsupervised method to align movie scripts with plot summaries using a global optimization model.
Outcome: The proposed method outperforms a baseline alignment model on ten movies with 76% F1 score.
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences.
Approach: They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality .
Outcome: The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research .
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language.
Approach: They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Outcome: The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language.
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that multimodal, multimodal approaches to lyrics translation are more effective than text-only approaches.
Approach: They propose a multilingual, multimodal benchmark for singable lyrics translation . they propose syllable-constrained audio-video LLM with Chain-of-Thought .
Outcome: The proposed system outperforms text-based models in singability and contextual accuracy.
MovieUN: A Dataset for Movie Understanding and Narrating (2022.findings-emnlp)

Copied to clipboard

Challenge: Automatic movie narration generation and narration grounding are important to provide a true movie experience for the blind and visually impaired.
Approach: They propose to use movie clips as a benchmark to support automatic movie narration generation and narration grounding tasks.
Outcome: The proposed methods are effective in supporting two movie-based tasks for the blind and visually impaired.
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to visual-language understanding lack unified tokenization for images and videos . lack of unified visual representations makes it difficult to learn multi-modal interactions from poor projection layers.
Approach: They propose to unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.
Outcome: The proposed model outperforms Video-ChatGPT on image benchmarks and on 9 image benchmark benchmarks.
Multi-view Story Characterization from Movie Plot Synopses and Reviews (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for characterizing stories by generating tags from synopses suffer from coverage issues.
Approach: They propose to use synopses and reviews to characterize stories by inferring attributes such as theme and style from written synopsis and reviews.
Outcome: The proposed model improves over methods that only use synopses and reviews . it can extract a complementary set of story attributes from reviews without supervision .
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on dense video captioning and video story generation have made some progress, but in practical applications, we typically require synchronized narrations for ongoing visual scenes.
Approach: They propose a task of Synchronized Video Storytelling to generate synchronized narrations for videos using a benchmark dataset with rich annotations.
Outcome: The proposed framework can generate narrations with the guidance of the generated or predefined storyline and human evaluations validate the effectiveness.
MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multilingual anaphora resolution include images and video inputs.
Approach: They propose to include multimodal information in the form of images in anaphora resolution tasks.
Outcome: The proposed approach improves resolution by 10% for unseen languages.
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)

Copied to clipboard

Challenge: a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations .
Approach: They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search .
Outcome: The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations