Speaker Naming in Movies (N18-1)

Copied to clipboard

Challenge: Identifying speakers and their names in movies is a primary task for many video analysis problems, such as automatic subtitle labeling.
Approach: They propose a model that leverages visual, textual, and acoustic modalities in an unified optimization framework for speaker naming in movies.
Outcome: The proposed model outperforms baseline models on the MovieQA 2017 challenge for speaker naming in movies and TV shows on visual, textual, and acoustic modalities.

Similar Papers

Towards Neural Speaker Modeling in Multi-Party Conversation: The Task, Dataset, and Models (L18-1)

Copied to clipboard

Challenge: Existing methods for speaker modeling are based on hand-crafted statistics and ad hoc to a certain application.
Approach: They propose to use speaker classification as a surrogate task for general speaker modeling and collect massive data to facilitate research in this direction.
Outcome: The proposed models outperform the existing models and are feasible with speaker identity information.
Reducing Sensitivity on Speaker Names for Text Generation from Dialogues (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models are sensitive to nuances, resulting in unfairness in real-world applications.
Approach: They propose to quantitatively measure a model's sensitivity on speaker names and comprehensively evaluate a number of known methods for reducing speaker name sensitivity.
Outcome: The proposed approach reduces speaker name sensitivity and improves quality of generation.
AligNarr: Aligning Narratives on Movies (2021.acl-short)

Copied to clipboard

Challenge: Experimental results show the viability of an unsupervised approach to align movie scripts with plot summaries.
Approach: They propose an unsupervised method to align movie scripts with plot summaries using a global optimization model.
Outcome: The proposed method outperforms a baseline alignment model on ten movies with 76% F1 score.
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 TV and movie series is filled with challenging multi-party dialogues.
Approach: They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues.
Outcome: The proposed dataset is a step towards better multi-party dialogue structuring and understanding.
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences.
Approach: They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality .
Outcome: The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research .
Serial Speakers: a Dataset of TV Series (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset of 155 episodes from popular american TV series is available to researchers . the dataset includes annotations for every speech turn (boundaries, speaker) and scene boundary .
Approach: They provide annotated dataset of 155 episodes from three popular american TV serials . they publicly release annotations for every speech turn (boundaries, speaker) and scene boundary .
Outcome: The dataset includes 155 episodes from three popular american TV serials: “Breaking Bad”, “Game of Thrones” and “House of Cards”.
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays (2024.findings-acl)

Copied to clipboard

Challenge: Movie screenplay summarization requires an understanding of long input contexts and elements unique to movies.
Approach: They propose a dataset for movie screenplay summarization that includes movie screenplayers accompanied by their Wikipedia plot summaries.
Outcome: The proposed dataset includes 2200 movie screenplays accompanied by their Wikipedia plot summaries.
Multimodal Large Language Models for Text-rich Image Understanding: A Comprehensive Review (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in vision-language models have unified perception and understanding tasks within Visual Question Answering paradigms.
Approach: They propose to outline timeline, architecture, and pipeline of nearly all TIU MLLMs and review their performance on mainstream benchmarks.
Outcome: The proposed models perform well on mainstream benchmarks and are compared with other models.
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation.
Approach: They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks.
Outcome: The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results.
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks.
Approach: They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications .
Outcome: The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations