Challenge: Existing methods for temporal sentence localization in videos focus on visual content, but they are insufficient to model complex video contents.
Approach: They propose a deep rectification-modulation network to correct attention misalignment . they use sentence information to capture frame-to-frame relation .
Outcome: The proposed method achieves state-of-the-art performance on three public datasets.

Similar Papers

On Pursuit of Designing Multi-modal Transformer for Video Grounding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for video grounding are not end-to-end, i.e., they rely on time-consuming post-processing steps to refine predictions.
Approach: They propose an end-to-end multi-modal Transformer model that uses two encoders and a cross-modal decoder for grounding prediction.
Outcome: The proposed model is 4.9% faster than existing models and is based on a set of encodings and decoders.
Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning (N18-2)

Copied to clipboard

Challenge: Existing multi-modal fusion methods have shown encouraging results in video understanding, but how to selectively fuse the multi-dimensional representations at different levels of details remains unexplored.
Approach: They propose a hierarchically aligned cross-modal attention framework to fuse audio and visual cues at different levels of detail.
Outcome: The proposed framework outperforms the previous best systems on the video captioning task.
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on temporal sentence grounding rely on expensive video-query paired annotations . despite this, there are no ground-truth annotations in the current work .
Approach: They propose to use paired video-query and segment boundary annotations to generate temporal sentence grounding without training.
Outcome: The proposed model outperforms existing unsupervised methods and beats supervised ones on two challenging datasets.
Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn effective alignment between vision and language features are insufficient in practice due to complicated multi-step reasoning.
Approach: They propose an iterative alignment network which iterates inter- and intra-modal features within multiple steps for more accurate grounding.
Outcome: The proposed model performs better than the state-of-the-arts on three challenging benchmarks.
Temporally Grounding Natural Sentence in Video (D18-1)

Copied to clipboard

Challenge: Existing methods for grounding natural sentences in video are limited to a single pass.
Approach: They propose a Temporal GroundNet (TGN) method that captures the evolving fine-grained frame-by-word interactions between video and sentence to ground the segment corresponding to the sentence.
Outcome: The proposed method significantly improves on the state-of-the-art methods on three public datasets and shows significant improvements in performance.
ExCL: Extractive Clip Localization Using Natural Language Descriptions (N19-1)

Copied to clipboard

Challenge: Prior approaches to retrieving clips within videos based on a given query are inefficient and text-clip similarity driven ranking-based approaches are far more complicated.
Approach: They propose an extractive approach that extracts the start and end frames by leveraging cross-modal interactions between the text and video to generate a joint representation.
Outcome: The proposed approach significantly outperforms state-of-the-art on two datasets and has comparable performance on a third.
Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video (P19-1)

Copied to clipboard

Challenge: Existing techniques for weakly-supervised spatio-temporally grounding natural sentence in video are lacking .
Approach: They propose a weakly-supervised task for spatially grounding sentences in video . they extract instances from video and encode them using attentive interactor . results demonstrate superiority of their proposed task over baseline approaches .
Outcome: The proposed model outperforms baseline approaches in a weakly-supervised task . it can characterize reliable instance-sentence pairs and penalize unreliable ones .
VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts.
Approach: They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels.
Outcome: The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels.
DeepMaven: Deep Question Answering on Long-Distance Movie/TV Show Videos with Multimedia Knowledge Extraction and Synthesis (2023.eacl-main)

Copied to clipboard

Challenge: Long video content understanding poses a challenging set of research questions as it involves long-distance, cross-media reasoning and knowledge awareness.
Approach: They propose a framework which extracts events, entities, and relations from the rich multimedia content in long videos to pre-construct movie knowledge graphs.
Outcome: The proposed framework performs competitively for both the new DeepMovieQA and the pre-existing MovieQA dataset.
Mitigating the Discrepancy Between Video and Text Temporal Sequences: A Time-Perception Enhanced Video Grounding method for LLM (2025.coling-main)

Copied to clipboard

Challenge: Existing video LLMs excel at capturing the overall description of a video but lack the ability to demonstrate an understanding of temporal dynamics and localized content within the video.
Approach: They propose a Time-Perception Enhanced Video Grounding via Boundary Perception and Temporal Reasoning to improve LLMs' understanding of video temporality.
Outcome: The proposed method improves on three datasets: ActivityNet, Charades, and DiDeMo (up to 11.2% improvement on R@0.3).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations