Challenge: Existing summarisation methods take no account of multimodal high-level paralinguistic features which form part of audio-visual presentations.
Approach: They propose to use audiovisual recordings to extract paralinguistic features from audio recordings . they use manual annotations to help users find relevant material .
Outcome: The proposed method can identify the most important or emphasised material within a presentation.

Similar Papers

MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities.
Approach: They propose a multimodal article and video summarization dataset that integrates resources from different modalities.
Outcome: The proposed dataset validates the important assistance role of external information for multimodal summarization.
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community .
Approach: They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations.
Outcome: The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community .
Large Scale Multi-Lingual Multi-Modal Summarization Dataset (2023.eacl-main)

Copied to clipboard

Challenge: a large dataset of document-image pairs and annotated multi-modal summarization data is needed for multi-lingual modeling . encoder-decoder models represent information comprising multiple modalities.
Approach: They propose to use a multi-lingual summarization dataset to analyze multi-modal summarizing using multi-linguistic annotated data.
Outcome: The proposed dataset is the largest multi-lingual multi-modal summarization dataset for 13 languages and consists of cross-lingual summarizing data for 2 languages.
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations (2025.acl-long)

Copied to clipboard

Challenge: VISTA dataset contains 18,599 recorded AI conference presentations . large multimodal models exhibit reduced performance in scientific contexts, study shows .
Approach: They propose a dataset specifically designed for video-to-text summarization in scientific domains.
Outcome: This paper compares the performance of large models with human models and shows that they improve on human models.
DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on multi-party dialogue discourse parsing focus on textual modality and two-party dialog . et al., 2016) focused on text-based discourse parses, ignoring the complexity and richness of multimodal interactions in real-world scenarios.
Approach: They construct the first publicly available English multimodal dataset for multi-party dialogue discourse parsing based on American TV dramas.
Outcome: The proposed dataset contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, covering rich multi-party interaction scenarios.
Bazinga! A Dataset for Multi-Party Dialogues Structuring (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 TV and movie series is filled with challenging multi-party dialogues.
Approach: They propose a dataset built around 16 TV and movie series with challenging multi-party dialogues.
Outcome: The proposed dataset is a step towards better multi-party dialogue structuring and understanding.
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language.
Approach: They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Outcome: The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language.
MMAR: Multilingual and Multimodal Anaphora Resolution in Instructional Videos (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to multilingual anaphora resolution include images and video inputs.
Approach: They propose to include multimodal information in the form of images in anaphora resolution tasks.
Outcome: The proposed approach improves resolution by 10% for unseen languages.
A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to summarize video content have only considered video and image data, and the trend towards multimodal video summarization is changing.
Approach: They propose a multimodal video summarization task setting and a dataset to train and evaluate the task.
Outcome: The proposed task is useful as a practical application and presents a highly challenging problem worthy of study.
M3AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset (2024.acl-long)

Copied to clipboard

Challenge: Publishing open-source academic video recordings is an emerging approach to sharing knowledge online.
Approach: They propose a multimodal, multigenre, and multipurpose audio-visual academic lecture dataset with human annotations for multimodal content recognition and understanding tasks.
Outcome: The proposed dataset can be used for multiple audio-visual recognition and understanding tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations