End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video Grounding (2022.acl-long)
Copied to clipboard
Mengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang, Zhou Zhao, Jiaxu Miao, Wenqiao Zhang, Wenming Tan, Jin Wang, Peng Wang, Shiliang Pu, Fei Wu
| Challenge: | Existing methods for grounding video frames with dense annotations require enormous amount of human effort. |
| Approach: | They propose to ground natural language in video frames with only one frame labeled . they propose an end-to-end model that eliminates interference of irrelevant frames . |
| Outcome: | The proposed model can ground natural language in all video frames with only one frame labeled . the proposed model eliminates interference of irrelevant frames based on branch search and cropping techniques . |
Similar Papers
DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization (D19-1)
Copied to clipboard
| Challenge: | Existing models for natural language video localization are top-down and bottom-up . however, both approaches suffer several limitations, leading to performance degradation . |
| Approach: | They propose a top-down approach for localizing a natural language description in a video sequence . they propose 'DEnse Bottom-Up Grounding' which uses the temporal boundaries of each video frame . |
| Outcome: | The proposed framework matches the speed of top-down models while surpassing the state-of-the-art models. |
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models (2025.findings-emnlp)
Copied to clipboard
Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao, Yufan Zhou, Yixin Cao, Qifan Wang, Weifeng Ge, Lifu Huang
| Challenge: | Video Large Language Models (VLMs) have been praised for their performance in coarse-grained video understanding but still face ineffective temporal grounding and inadequate timestamp representations. |
| Approach: | They propose a novel Video-LLM that senses and reasoned over specific video moments with fine-grained temporal precision. |
| Outcome: | The proposed model surpasses existing models in fine-grained video understanding tasks and exhibits strong potential as a general video understanding assistant. |
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Long video understanding presents unique challenges due to the complexity of reasoning over extended timespans. |
| Approach: | They propose a framework VideoINSTA to leverage large language models for video understanding . they propose 'event-based temporalreasoning' and 'content-based spatial reasoning' |
| Outcome: | The proposed model significantly improves state-of-the-art on three long video question-answering benchmarks. |
On Pursuit of Designing Multi-modal Transformer for Video Grounding (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for video grounding are not end-to-end, i.e., they rely on time-consuming post-processing steps to refine predictions. |
| Approach: | They propose an end-to-end multi-modal Transformer model that uses two encoders and a cross-modal decoder for grounding prediction. |
| Outcome: | The proposed model is 4.9% faster than existing models and is based on a set of encodings and decoders. |
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction (2021.naacl-main)
Copied to clipboard
| Challenge: | Scholarly work in this area uses toy worlds and synthetic linguistic data, but grounded language learning offers several practical and scientific advantages. |
| Approach: | They propose to model teacher-learner dynamics through natural interactions occurring between users and search engines. |
| Outcome: | The proposed model is better than non-grounded models on compositionality and zero-shot inference tasks. |
Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search (P18-1)
Copied to clipboard
| Challenge: | a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search is currently used to learn word representations. |
| Approach: | They propose a large-scale lookup operation to ground language via ‘snapshots’ of our physical world accessed through image search. |
| Outcome: | The proposed model is based on a large-scale lookup operation to ground language using image search. |
VLURes: Benchmarking Long-Text Grounding and Cross-Lingual Robustness in Vision Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | ***VLURes** provides a practical testbed for long-text grounding and multilingual robustness in web-realistic agent settings. |
| Approach: | They propose a multilingual benchmark for evaluating vision-language models under long-text grounding. |
| Outcome: | ***VLURes** provides a testbed for long-text grounding and multilingual robustness in web-realistic agent settings. |
Towards Parameter-Efficient Integration of Pre-Trained Language Models In Temporal Video Grounding (2023.findings-acl)
Copied to clipboard
Erica Kido Shimomoto, Edison Marrese-Taylor, Hiroya Takamura, Ichiro Kobayashi, Hideki Nakayama, Yusuke Miyao
| Challenge: | Recent studies have improved query inputs with pre-trained language models, but the effects of this integration are unclear. |
| Approach: | They propose to integrate query sentences with pre-trained language models to train TVG models. |
| Outcome: | The proposed model integrates query sentences with pre-trained language models at cost of more expensive training. |
Read Before Grounding: Scene Knowledge Visual Grounding via Multi-step Parsing (2025.coling-main)
Copied to clipboard
| Challenge: | Existing VG datasets use simple textual descriptions with limited attribute and spatial information between images and text. |
| Approach: | They propose a method that transforms visual knowledge into concise, information-dense visual descriptions. |
| Outcome: | The proposed method significantly improves performance of multimodal grounding models. |
Grounding language acquisition by training semantic parsers using captioned videos (D18-1)
Copied to clipboard
| Challenge: | a new method for parsing sentences using captioned videos is being developed . we use video clips to ground the semantics of language, but without annotations . |
| Approach: | They develop a semantic parser that is trained in a grounded setting using captioned videos . they use a corpus of sentences paired with videos without other annotations to train it . |
| Outcome: | The proposed parser recovers the meaning of English sentences despite no annotations . learning a grounded semantic parsers can expand the range of data that parseurs can be trained on . |