Multilingual Synopses of Movie Narratives: A Dataset for Vision-Language Story Understanding (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Story video-text alignment is a core task in computational story understanding, but its progress has been held back by the scarcity of manually annotated video- text correspondences and the heavy concentration on English narrations of Hollywood movies. |
| Approach: | They construct a multilingual video story dataset with 13,166 movie summary videos from 7 languages and manual annotations of fine-grained video-text correspondences. |
| Outcome: | The proposed approach outperforms the SOTA methods on clip accuracy and Sentence IoU scores. |
Similar Papers
AligNarr: Aligning Narratives on Movies (2021.acl-short)
Copied to clipboard
| Challenge: | Experimental results show the viability of an unsupervised approach to align movie scripts with plot summaries. |
| Approach: | They propose an unsupervised method to align movie scripts with plot summaries using a global optimization model. |
| Outcome: | The proposed method outperforms a baseline alignment model on ten movies with 76% F1 score. |
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)
Copied to clipboard
| Challenge: | Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. |
| Approach: | They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality . |
| Outcome: | The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research . |
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language. |
| Approach: | They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. |
| Outcome: | The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language. |
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Experimental results show that multimodal, multimodal approaches to lyrics translation are more effective than text-only approaches. |
| Approach: | They propose a multilingual, multimodal benchmark for singable lyrics translation . they propose syllable-constrained audio-video LLM with Chain-of-Thought . |
| Outcome: | The proposed system outperforms text-based models in singability and contextual accuracy. |
MovieUN: A Dataset for Movie Understanding and Narrating (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Automatic movie narration generation and narration grounding are important to provide a true movie experience for the blind and visually impaired. |
| Approach: | They propose to use movie clips as a benchmark to support automatic movie narration generation and narration grounding tasks. |
| Outcome: | The proposed methods are effective in supporting two movie-based tasks for the blind and visually impaired. |
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to visual-language understanding lack unified tokenization for images and videos . lack of unified visual representations makes it difficult to learn multi-modal interactions from poor projection layers. |
| Approach: | They propose to unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM. |
| Outcome: | The proposed model outperforms Video-ChatGPT on image benchmarks and on 9 image benchmark benchmarks. |
Multi-view Story Characterization from Movie Plot Synopses and Reviews (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for characterizing stories by generating tags from synopses suffer from coverage issues. |
| Approach: | They propose to use synopses and reviews to characterize stories by inferring attributes such as theme and style from written synopsis and reviews. |
| Outcome: | The proposed model improves over methods that only use synopses and reviews . it can extract a complementary set of story attributes from reviews without supervision . |
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies on dense video captioning and video story generation have made some progress, but in practical applications, we typically require synchronized narrations for ongoing visual scenes. |
| Approach: | They propose a task of Synchronized Video Storytelling to generate synchronized narrations for videos using a benchmark dataset with rich annotations. |
| Outcome: | The proposed framework can generate narrations with the guidance of the generated or predefined storyline and human evaluations validate the effectiveness. |
MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to multilingual anaphora resolution include images and video inputs. |
| Approach: | They propose to include multimodal information in the form of images in anaphora resolution tasks. |
| Outcome: | The proposed approach improves resolution by 10% for unseen languages. |
Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models (2021.naacl-main)
Copied to clipboard
| Challenge: | a new study examines zero-shot cross-lingual transfer of vision-language models . we study multilingual text-to-video search in non-English languages without annotations . |
| Approach: | They propose a Transformer-based model that learns contextual multilingual multimodal embeddings . they propose 'zero-shot cross-lingual transfer' to improve multilingual search . |
| Outcome: | The proposed model outperforms baselines on multilingual text-to-video search and multilingual image search on VTT and VATEX. |