Challenge: Prior approaches to retrieving clips within videos based on a given query are inefficient and text-clip similarity driven ranking-based approaches are far more complicated.
Approach: They propose an extractive approach that extracts the start and end frames by leveraging cross-modal interactions between the text and video to generate a joint representation.
Outcome: The proposed approach significantly outperforms state-of-the-art on two datasets and has comparable performance on a third.

Similar Papers

Reasoning Step-by-Step: Temporal Sentence Localization in Videos via Deep Rectification-Modulation Network (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for temporal sentence localization in videos focus on visual content, but they are insufficient to model complex video contents.
Approach: They propose a deep rectification-modulation network to correct attention misalignment . they use sentence information to capture frame-to-frame relation .
Outcome: The proposed method achieves state-of-the-art performance on three public datasets.
Exploiting Discourse-Level Segmentation for Extractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing approaches to extract summarize text are based on sentences as the elementary unit, but semantic segments containing supplementary information or descriptive details are often nonessential in the generated summaries.
Approach: They propose to exploit discourse-level segmentation as a finer-grained means to more precisely pinpoint the core content in a document.
Outcome: The proposed method improves extractive summarization performance on CNN/Daily Mail dataset.
Cross-modal Contrastive Learning for Speech Translation (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches for speech translation focus on using additional data from MT and automatic speech recognition (ASR).
Approach: They propose a cross-modal contrastive learning method for end-to-end speech-totext translation.
Outcome: The proposed method outperforms existing methods on a popular benchmark MuST-C.
On Pursuit of Designing Multi-modal Transformer for Video Grounding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for video grounding are not end-to-end, i.e., they rely on time-consuming post-processing steps to refine predictions.
Approach: They propose an end-to-end multi-modal Transformer model that uses two encoders and a cross-modal decoder for grounding prediction.
Outcome: The proposed model is 4.9% faster than existing models and is based on a set of encodings and decoders.
Retrieve Fast, Rerank Smart: Cooperative and Joint Approaches for Improved Cross-Modal Retrieval (2022.tacl-1)

Copied to clipboard

Challenge: Current approaches to cross-modal retrieval process text and visual input jointly . current approaches are pretrained from scratch and suffer from huge retrieval latency and inefficiency issues .
Approach: They propose a cooperative retrieve-and-rerank framework that turns pretrained text-image multi-modal models into efficient retrieval models.
Outcome: The proposed framework improves retrieval performance over current approaches . it uses twin networks to encode all items of a corpus and a cross-encoder component for a more nuanced ranking .
End-to-end Knowledge Retrieval with Multi-modal Queries (2023.acl-long)

Copied to clipboard

Challenge: a new task is proposed to learn knowledge retrieval with multimodal queries . a vision-language model can retrieve knowledge using images and text inputs .
Approach: They propose a task for vision-language models to retrieve knowledge with multi-modal queries . they propose reViz, a model that integrates content from both text and image queries based on a multimodal query task .
Outcome: The proposed task performs better under zero-shot settings than previous work on cross-modal retrieval.
Neural Sequence Segmentation as Determining the Leftmost Segments (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to segment sentences are mostly at token level, limiting their full potential to capture long-term dependencies.
Approach: They propose a framework that incrementally segments natural language sentences at segment level.
Outcome: The proposed framework outperforms baseline methods on syntactic chunking and Chinese part-of-speech tagging datasets.
Extending CLIP’s Image-Text Alignment to Referring Image Segmentation (2024.naacl-long)

Copied to clipboard

Challenge: Referring Image Segmentation (RIS) is a cross-modal task that aims to segment an instance described by a natural language expression.
Approach: They propose a framework that leverages the cross-modal nature of CLIP for RIS by leveraging image-text alignment knowledge in CLIP's image-embedding space.
Outcome: The proposed framework outperforms CLIP-based methods on all three major RIS benchmarks and outperformed previous CLIP methods.
Natural Language Video Localization with Learnable Moment Proposals (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for video moment localization have poor performance due to predefined rules.
Approach: They propose a model with a fixed set of learnable moment proposals with 'border-aware loss' they propose to localize the video moment corresponding to the query by locating the start and end timestamps in an untrimmed video.
Outcome: The proposed model outperforms state-of-the-art models on two challenging benchmarks.
On the Language Encoder of Contrastive Cross-modal Models (2024.findings-acl)

Copied to clipboard

Challenge: Pretrained audio-language models such as AudioCLIP and AudioCLAP have shown promising results on vision-language (VL) tasks.
Approach: They extensively evaluate how unsupervised and supervised sentence embedding training affect language encoder quality and cross-modal task performance.
Outcome: The proposed model improves on visual-language (VL) and audio-language tasks when the amount of training data is large.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations