MMAD:Multi-modal Movie Audio Description (2024.lrec-main)

Copied to clipboard

Challenge: Current methods of creating accessible movies rely on manual work, resulting in high costs and limited scalability.
Approach: They propose a multi-modal movie audio description pipeline that generates narrations of information that is not accessible through unimodal hearing in movies.
Outcome: The proposed pipeline surpasses existing baselines in performance on widely used datasets.

Similar Papers

Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies (2025.findings-naacl)

Copied to clipboard

Challenge: Audio descriptions (ADs) are acoustic commentaries designed to assist blind and visually impaired individuals in accessing digital media content.
Approach: They examine how state-of-the-art NLP and CV technologies can be applied to generate ADs . they identify essential research directions for the future .
Outcome: The proposed technologies can be applied to generate audio descriptions (ADs) the process is time-consuming and costly, and requires significant human effort . the authors identify key research directions for the future .
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information.
Approach: They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions .
Outcome: The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets.
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences.
Approach: They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality .
Outcome: The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research .
What You See is What You Ask: Evaluating Audio Descriptions (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies evaluate audio descriptions (ADs) using trimmed clips, but writing them is subjective.
Approach: They propose a QA benchmark that evaluates audio descriptions at the level of short, coherent video segments.
Outcome: The proposed evaluation paradigm addresses two themes central to ADs . it compares two humannarrated AD tracks and shows that current methods lag behind human-authored ADs.
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding (2023.emnlp-demo)

Copied to clipboard

Challenge: Large Language Models (LLMs) are capable of understanding multi-modal content, but textonly human-computer interaction is not sufficient for many application scenarios.
Approach: They propose a video-to-text generation task and a multi-modal framework that bootstraps cross-modal training from frozen pre-trained visual & audio encoders and frozen LLMs.
Outcome: The proposed framework can understand both visual and auditory content in video and generate meaningful responses grounded in the visual and audio information presented in the videos.
Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio Captioning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing AAC datasets suffer from short and simplistic captions, limiting expressiveness and semantic depth.
Approach: They propose a multi-modal dataset that pairs audio with corresponding video and leverages large language models to generate rich, descriptive captions.
Outcome: The proposed framework outperforms existing benchmarks in caption length, lexical diversity, and human-rated quality.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed .
Approach: They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities.
Outcome: The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech .
MovieUN: A Dataset for Movie Understanding and Narrating (2022.findings-emnlp)

Copied to clipboard

Challenge: Automatic movie narration generation and narration grounding are important to provide a true movie experience for the blind and visually impaired.
Approach: They propose to use movie clips as a benchmark to support automatic movie narration generation and narration grounding tasks.
Outcome: The proposed methods are effective in supporting two movie-based tasks for the blind and visually impaired.
Sound of Story: Multi-modal Storytelling with Audio (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on storytelling with sound have focused on visuals and sounds, but little attention has been given to sound.
Approach: They propose to establish a new component called background sound which is story context-based audio without any linguistic information.
Outcome: The proposed dataset is the largest well-curated dataset for storytelling with sound . it contains 27,354 stories with 19.6 images per story and 984 hours of speech-decoupled audio .
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization (2024.lrec-main)

Copied to clipboard

Challenge: Existing product summarization methods lack end-to-end product summaries and multi-grained multi-modal modeling.
Approach: They propose an end-to-end multi-grained multi-modal attribute-aware product summarization method that jointly models product attributes and generates product summaries.
Outcome: The proposed method outperforms state-of-the-art product summarization methods on a large-scale Chinese e-commence dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations