Challenge: Multimodal semantic comprehension has attracted increasing research interest recently such as visual question answering and caption generation.
Approach: They propose to use a large-scale multimodal instructional video dataset to support fine-grained comprehension research in specific domain.
Outcome: The proposed dataset contains 2,800 videos from YouTube, spanning more than 420 hours in total.

Similar Papers

TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing text-rich image understanding benchmarks lack scale and fragmented scenarios . a new full-image structured output format is proposed to enable fine-grained evaluation of perception and reasoning capabilities.
Approach: They propose a large-scale, multilingual benchmark that includes over 100,000 annotations and 22,000 question-answer pairs.
Outcome: The proposed framework provides a comprehensive platform for developing and evaluating next-generation multimodal AI systems.
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information.
Approach: They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions .
Outcome: The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets.
mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with Images (2023.emnlp-main)

Copied to clipboard

Challenge: Existing summarization datasets do not cover multimodal discussions, multiple modalities, or both . mRedditSum consists of 3,033 discussion threads and images with human-written summaries.
Approach: They propose a multimodal discussion summarization dataset that annotates 3,033 discussion threads with a human-written summary.
Outcome: The proposed method outperforms existing models and serves as competitive baseline for future work.
Multimodal Pretraining for Dense Video Captioning (2020.aacl-main)

Copied to clipboard

Challenge: a billion hours of videos are being watched on YouTube every day . videos are difficult to skim through, making it harder to quickly target the relevant part(s) of a video.
Approach: They propose to use a video timeline tag dataset to generate time-stamped annotations for videos . they propose to pretrain and finetune captioning models using YouCook2 and ViTT .
Outcome: The proposed model generalizes well and is robust over a wide variety of instructional videos.
ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos (2023.emnlp-main)

Copied to clipboard

Challenge: despite its importance, there are few datasets that cover multimodal counterfactual reasoning . a dataset focusing on this area is limited because of its limited coverage over synthetic environments .
Approach: They develop a video question answering dataset that provides questions on multimodal reasoning . they ask questions about counterfactual hypotheses over visual events .
Outcome: The proposed dataset shows a significant performance gap between models and humans . it provides questions that span physical, social, and temporal dimensions .
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities.
Approach: They propose a multimodal article and video summarization dataset that integrates resources from different modalities.
Outcome: The proposed dataset validates the important assistance role of external information for multimodal summarization.
MultiSkill: Evaluating Large Multimodal Models for Fine-grained Alignment Skills (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation settings for large multimodal models focus on coarse-grained evaluation without considering skill composition required by specific instructions.
Approach: They propose an evaluation protocol that assesses large multimodal models across multiple fine-grained skills for alignment with human values.
Outcome: The proposed evaluation protocol decomposes coarse-level scoring to fine-grained skill set-level score tailored to each instruction.
Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models that understand image and text but also cross-reference in-between are lacking in evaluation data resources.
Approach: They propose a multimodal evaluation pipeline to automatically generate question-answer pairs to test models’ understanding of the visual scene, text, and related knowledge.
Outcome: The proposed model can answer the highly semantic VCR question correctly but fails to answer related visual question (Q2), textual question (q3), and background knowledge question ( Q4) as shallow mappings with language priors and unbalanced utilization of information between modalities.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in text-video retrieval neglect the crucial user perspective, leading to discrepancies between user queries and content retrieved.
Approach: They propose a novel, training-free pipeline that leverages Large Language Models for iterative feedback learning.
Outcome: Experimental results show that MERLIN significantly outperforms existing systems in video retrieval.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations