| Challenge: | Abstractive summarization is a task of producing a shorter version of the content in the document while preserving its information. |
| Approach: | They propose a new evaluation metric that measures semantic adequacy rather than fluency of abstractive summarization tasks. |
| Outcome: | The proposed model integrates information from different sources into a coherent output. |
Similar Papers
Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing multimodal summarization methods are limited to monolingual videos . a proposed task aims to generate cross-lingual summaries from multimodal inputs . |
| Approach: | They propose a task to generate cross-lingual summaries from multimodal inputs of videos . they propose fusion network that integrates multimodal and cross-linguistic information . |
| Outcome: | The proposed task outperforms existing methods on a reorganized How2 dataset on the reorganized How2 data set. |
UrduMASD: A Multimodal Abstractive Summarization Dataset for Urdu (2024.lrec-main)
Copied to clipboard
| Challenge: | a surge of multimodal content on social media has transformed our methods of communication and information exchange. |
| Approach: | They propose a video-based Urdu multimodal abstractive text summarization dataset . it uses a variety of evaluation metrics to ensure the quality of the dataset amounted to a high quality one . |
| Outcome: | The proposed dataset surpasses existing datasets on key quality metrics. |
Hierarchical3D Adapters for Long Video-to-text Summarization (2023.findings-eacl)
Copied to clipboard
| Challenge: | a recent study shows that multimodal summarization is not efficient for long inputs and outputs. |
| Approach: | They extend a TV episode transcript summarization dataset and create a multimodal variant by collecting full-length videos. |
| Outcome: | The proposed model can be tuned to perform multimodal summarization tasks efficiently using adapter modules augmented with a hierarchical structure while tuning only 3.8% of model parameters. |
A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to summarize video content have only considered video and image data, and the trend towards multimodal video summarization is changing. |
| Approach: | They propose a multimodal video summarization task setting and a dataset to train and evaluate the task. |
| Outcome: | The proposed task is useful as a practical application and presents a highly challenging problem worthy of study. |
Abstractive Multi-Video Captioning: Benchmark Dataset Construction and Extensive Evaluation (2024.lrec-main)
Copied to clipboard
| Challenge: | Abstractive multi-video captioning focuses on abstracting multiple videos with natural language. |
| Approach: | They propose a task that generates an abstract caption of shared video content . they propose end-to-end and cascade approaches to abstractive multi-video captioning . |
| Outcome: | The proposed task generates an abstract caption of shared content in a video group containing multiple videos. |
VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies show that multimodal news can significantly improve users' sense of satisfaction for informativeness. |
| Approach: | They propose a task of Video-based Multimodal Summarization with Multimodal Output to solve this problem. |
| Outcome: | The proposed method can generate multimodal summaries with a single input . it can model the temporal dependency of video with semantic meaning of article . |
What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific Presentations (2025.acl-long)
Copied to clipboard
Dongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon, Rohit Saxena, Zheng Zhao, Yifu Qiu, Mirella Lapata, Vera Demberg
| Challenge: | VISTA dataset contains 18,599 recorded AI conference presentations . large multimodal models exhibit reduced performance in scientific contexts, study shows . |
| Approach: | They propose a dataset specifically designed for video-to-text summarization in scientific domains. |
| Outcome: | This paper compares the performance of large models with human models and shows that they improve on human models. |
Summary-Oriented Vision Modeling for Multimodal Abstractive Summarization (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies on multimodal abstractive summarization focus on how to use extracted visual features to produce a concise summary given the multimodal data. |
| Approach: | They propose to improve the visual quality of the multimodal abstractive summarization model by capturing summary-oriented visual features. |
| Outcome: | The proposed approach achieves state-of-the-art under 44 languages and is highly effective on high-resource English datasets. |
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities. |
| Approach: | They propose a multimodal article and video summarization dataset that integrates resources from different modalities. |
| Outcome: | The proposed dataset validates the important assistance role of external information for multimodal summarization. |
Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting Summarization (P19-1)
Copied to clipboard
| Challenge: | Existing models for extractive summarization of meetings are unfocused and lack content coverage. |
| Approach: | They propose a multi-modal hierarchical attention model that prioritizes segmentation and summarization . they propose to use multi-level hierarchies to narrow down the focus into topically-relevant segments . |
| Outcome: | The proposed model outperforms the state-of-the-art with BLEU and ROUGE measures. |