| Challenge: | Identifying speakers and their names in movies is a primary task for many video analysis problems, such as automatic subtitle labeling. |
| Approach: | They propose a model that leverages visual, textual, and acoustic modalities in an unified optimization framework for speaker naming in movies. |
| Outcome: | The proposed model outperforms baseline models on the MovieQA 2017 challenge for speaker naming in movies and TV shows on visual, textual, and acoustic modalities. |
Similar Papers
Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models (L18-1)
Copied to clipboard
| Challenge: | Existing methods for speaker modeling are based on hand-crafted statistics and ad hoc to a certain application. |
| Approach: | They propose to use speaker classification as a surrogate task for general speaker modeling and collect massive data to facilitate research in this direction. |
| Outcome: | The proposed models outperform the existing models and are feasible with speaker identity information. |
Reducing Sensitivity on Speaker Names for Text Generation from Dialogues (2023.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained language models are sensitive to nuances, resulting in unfairness in real-world applications. |
| Approach: | They propose to quantitatively measure a model's sensitivity on speaker names and comprehensively evaluate a number of known methods for reducing speaker name sensitivity. |
| Outcome: | The proposed approach reduces speaker name sensitivity and improves quality of generation. |
AligNarr: Aligning Narratives on Movies (2021.acl-short)
Copied to clipboard
| Challenge: | Experimental results show the viability of an unsupervised approach to align movie scripts with plot summaries. |
| Approach: | They propose an unsupervised method to align movie scripts with plot summaries using a global optimization model. |
| Outcome: | The proposed method outperforms a baseline alignment model on ten movies with 76% F1 score. |
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)
Copied to clipboard
Paul Lerner, Juliette Bergoënd, Camille Guinaudeau, Hervé Bredin, Benjamin Maurice, Sharleyne Lefevre, Martin Bouteiller, Aman Berhe, Léo Galmant, Ruiqing Yin, Claude Barras
| Challenge: | a dataset of 16 TV and movie series is filled with challenging multi-party dialogues. |
| Approach: | They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues. |
| Outcome: | The proposed dataset is a step towards better multi-party dialogue structuring and understanding. |
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)
Copied to clipboard
| Challenge: | Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences. |
| Approach: | They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality . |
| Outcome: | The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research . |
Serial Speakers: a Dataset of TV Series (2020.lrec-1)
Copied to clipboard
| Challenge: | a new dataset of 155 episodes from popular american TV series is available to researchers . the dataset includes annotations for every speech turn (boundaries, speaker) and scene boundary . |
| Approach: | They provide annotated dataset of 155 episodes from three popular american TV serials . they publicly release annotations for every speech turn (boundaries, speaker) and scene boundary . |
| Outcome: | The dataset includes 155 episodes from three popular american TV serials: “Breaking Bad”, “Game of Thrones” and “House of Cards”. |
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays (2024.findings-acl)
Copied to clipboard
| Challenge: | Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies. |
| Approach: | They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries. |
| Outcome: | The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries. |
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)
Copied to clipboard
Pei Fu, Tongkun Guan, Zining Wang, Zhentao Guo, Chen Duan, Hao Sun, Boming Chen, Qianyi Jiang, Jiayao Ma, Kai Zhou, Junfeng Luo
| Challenge: | Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms. |
| Approach: | They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks. |
| Outcome: | The proposed models perform well on mainstream benchmarks and are compared with other models. |
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)
Copied to clipboard
Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski, Łukasz Bondaruk, Jakub Kubiak, Marcin Lewandowski, Marek Kubis
| Challenge: | Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation. |
| Approach: | They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks. |
| Outcome: | The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results. |
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks. |
| Approach: | They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications . |
| Outcome: | The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research. |