An Empirical Study of Frame Selection for Text-to-Video Retrieval (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for text-to-video retrieval select a subset of frames to represent video content . current methods only explore video contents while ignoring relevancy to texts . |
| Approach: | They propose to use a subset of frames to represent video content for TVR . they analyze six different frame selection methods to determine their effectiveness . |
| Outcome: | The proposed method improves retrieval efficiency without sacrificing visual details . the proposed method explores the video contents while ignoring relevancy to texts . |
Similar Papers
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter (2024.findings-acl)
Copied to clipboard
Meng Cao, Haoran Tang, Jinfa Huang, Peng Jin, Can Zhang, Ruyang Liu, Long Chen, Xiaodan Liang, Li Yuan, Ge Li
| Challenge: | Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. |
| Approach: | They propose to conduct efficient text-video Retrieval with a salient-and-correlated AdaPter . they propose a low-rank modulation module to refine per-image features from frozen CLIP backbone . |
| Outcome: | Experiments on four TVR datasets show that the proposed method performs better than other methods. |
VideoRAG: Retrieval-Augmented Generation over Video Corpus (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to generating models rely on text and images, but video content is a rich source of multimodal knowledge. |
| Approach: | They propose a framework that dynamically retrieves videos based on their relevance with queries . they use large video language models to represent video content for retrieval . |
| Outcome: | The proposed framework retrieves videos based on relevance with queries and integrates both visual and textual information. |
Self-Adaptive Sampling for Accurate Video Question Answering on Image Text Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Image–text models (ITMs) are the prevalent architecture to solve video question–answering tasks, which requires only a few input frames to save huge computational cost compared to video–language models. |
| Approach: | They propose a sampling method based on question–frame correlation that is efficient for the few-frame situations. |
| Outcome: | The proposed method can boost the performance of image–text pretrained models and have a wide application scenario in terms of model architectures and dataset types. |
Revealing Single Frame Bias for Video-and-Language Learning (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for video-and-language learning use multiple frames as inputs. |
| Approach: | They propose to use single-frame models for video-and-language learning to investigate temporality in video- and language tasks. |
| Outcome: | The proposed model does not take into account temporal information on video-and-language tasks. |
ELIOT: Zero-Shot Video-Text Retrieval through Relevance-Boosted Captioning and Structural Information Extraction (2025.naacl-srw)
Copied to clipboard
| Challenge: | Recent advances in video-text retrieval (VTR) have relied on supervised learning and fine-tuning. |
| Approach: | They propose a zero-shot video-text retrieval framework that leverages off-the-shelf captioners, large language models, and text retrieval methods without additional training or annotated data. |
| Outcome: | The proposed framework outperforms existing methods on video-text retrieval benchmarks without data. |
A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to summarize video content have only considered video and image data, and the trend towards multimodal video summarization is changing. |
| Approach: | They propose a multimodal video summarization task setting and a dataset to train and evaluate the task. |
| Outcome: | The proposed task is useful as a practical application and presents a highly challenging problem worthy of study. |
Frame2: A FrameNet-based Multimodal Dataset for Tackling Text-image Interactions in Video (2024.lrec-main)
Copied to clipboard
Frederico Belcavello, Tiago Timponi Torrent, Ely E. Matos, Adriana S. Pagano, Maucha Gamonal, Natalia Sigiliano, Lívia Vicente Dutra, Helen de Andrade Abreu, Mairon Samagaio, Mariane Carvalho, Franciany Campos, Gabrielly Azalim, Bruna Mazzei, Mateus Fonseca de Oliveira, Ana Carolina Loçasso Luz, Lívia Pádua Ruiz, Júlia Bellei, Amanda Pestana, Josiane Costa, Iasmin Rabelo, Anna Beatriz Silva, Raquel Roza, Mariana Souza, Igor Oliveira
| Challenge: | et al., 2016) describe a multimodal dataset built from a Brazilian travel TV show . frameNet is composed of frames and their associated roles in a network of typed frame-to-frame relations. |
| Approach: | They present a multimodal dataset built from a Brazilian travel TV show annotated for FrameNet categories for both text and image communicative modes. |
| Outcome: | The proposed dataset includes 230 minutes of video annotated for FrameNet categories . the model can be applied to other communicative modes, i.e., images . |
Contrastive Video-Language Learning with Fine-grained Frame Sampling (2022.aacl-main)
Copied to clipboard
| Challenge: | despite recent progress in video and language representation learning, the weak or sparse correspondence between the two modalities remains a bottleneck. |
| Approach: | They propose a fine-grained contrastive objective for video frame sampling to improve cross-modal correspondence. |
| Outcome: | The proposed approach achieves state-of-the-art performance on YouCookII with long videos. |
Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQA (2020.acl-main)
Copied to clipboard
| Challenge: | Recent years have witnessed a paradigm shift in the way we get our information, and a lot of it. |
| Approach: | They propose a video question answering model which integrates multi-modal input sources and finds temporally relevant information to answer questions. |
| Outcome: | The proposed model outperforms the state-of-the-art on a TVQA dataset. |
Fighting FIRe with FIRE: Assessing the Validity of Text-to-Video Retrieval Benchmarks (2023.findings-eacl)
Copied to clipboard
Pedro Rodriguez, Mahmoud Azab, Becka Silvert, Renato Sanchez, Linzy Labson, Hardik Shah, Seungwhan Moon
| Challenge: | Existing benchmarks for text-to-video retrieval are incomplete, resulting in false negatives . a recent state-of-the-art model gains 25% recall points, but this is not the case for TVR. |
| Approach: | They propose to retire video captioning datasets as TVR benchmarks . they propose to annotate and release additional caption-video pairs to mitigate this flaw . |
| Outcome: | The proposed method fails to accurately reflect reality, despite lack of purpose-built benchmarks. |