Papers with STA

12 papers
See More, Store Less: Memory-Efficient Resolution for Video Moment Retrieval (2026.findings-eacl)

Copied to clipboard

Challenge: Existing video moment retrieval methods rely on sparse frame sampling, risking information loss.
Approach: a new video-based framework enhances memory efficiency while maintaining high information resolution . SMORE uses query-guided captions to encode semantics aligned with user intent .
Outcome: a new framework improves memory efficiency while maintaining high information resolution . it achieves state-of-the-art performance on QVHighlights, Charades-STA, and ActivityNet-Captions benchmarks .
Exploiting Intrinsic Multilateral Logical Rules for Weakly Supervised Natural Language Video Localization (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for WS-NLVL rarely consider complex temporal relations enclosing the language query, yielding illogical predictions.
Approach: They propose a plug-and-play method to exploit temporal relations and logical rules for WS-NLVL.
Outcome: The proposed method is able to retrieve the moment corresponding to a language query in a video with only video-language pairs utilized during training.
Relation-aware Video Reading Comprehension for Temporal Language Grounding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for temporal language grounding in videos are boundary regression and span extraction tasks.
Approach: They propose a Relation-aware Network to localize a temporal span relevant to a given query sentence.
Outcome: The proposed framework selects a video moment choice from the predefined answer set with the aid of coarse-and-fine choice-query interaction and choice-choice relation construction.
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies on video large language models focus on model architectures and training datasets . interaction format between user and model is unsatisfactory for time-sensitive tasks .
Approach: They propose a video-text duet interaction format that allows for continuous playback of the video . when a text message ends, the video continues to play, similar to the alternative of two performers in a duet.
Outcome: The proposed format improves performance on time-sensitive tasks with minimal training efforts.
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing (2025.emnlp-main)

Copied to clipboard

Challenge: Current models struggle with long-form videos due to the quadratic complexity of attention mechanisms.
Approach: They propose a model-agnostic framework that leverages temporal cues from queries to prune video tokens.
Outcome: The proposed framework reduces computation by 65% while preserving 97-99% of original performance.
VideoQA-TA: Temporal-Aware Multi-Modal Video Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts.
Approach: They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels.
Outcome: The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels.
SA-DETR:Span Aware Detection Transformer for Moment Retrieval (2025.coling-main)

Copied to clipboard

Challenge: Moment Retrieval aims to locate video segments related to text.
Approach: They propose a method that leverages the importance of instance related span anchors . they initialize span anchor using instance related fuse token and supervise them with GT labels .
Outcome: The proposed method achieves competitive results on QVHighlights, Charades-STA and TACoS.
DEBUG: A Dense Bottom-Up Grounding Approach for Natural Language Video Localization (D19-1)

Copied to clipboard

Challenge: Existing models for natural language video localization are top-down and bottom-up . however, both approaches suffer several limitations, leading to performance degradation .
Approach: They propose a top-down approach for localizing a natural language description in a video sequence . they propose 'DEnse Bottom-Up Grounding' which uses the temporal boundaries of each video frame .
Outcome: The proposed framework matches the speed of top-down models while surpassing the state-of-the-art models.
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on temporal sentence grounding rely on expensive video-query paired annotations . despite this, there are no ground-truth annotations in the current work .
Approach: They propose to use paired video-query and segment boundary annotations to generate temporal sentence grounding without training.
Outcome: The proposed model outperforms existing unsupervised methods and beats supervised ones on two challenging datasets.
You Only Need One Single Token to Refine Safety Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Excessive safety can lead to over-refusal, where models reject harmful-looking yet benign queries, severely limiting utility.
Approach: They propose a lightweight training-based approach that reshapes the distributions of harmful and benign samples within the model’s decision space by using a single-token prefix.
Outcome: The proposed approach can distinguish between harmful and benign samples while keeping the model frozen.
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms (2025.acl-long)

Copied to clipboard

Challenge: Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering.
Approach: They propose a method that isolates and manipulates disentangled knowledge components to enhance safety by using sparse autoencoders to disentangle knowledge in high-dimensional spaces for steering.
Outcome: The proposed method is able to isolate and manipulate disentangled knowledge components to enhance safety in large reasoning models.
STA-CoT: Structured Target-Centric Agentic Chain-of-Thought for Consistent Multi-Image Geological Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Reliable multi-image geological reasoning is essential for automating expert tasks in remote-sensing mineral exploration.
Approach: They propose a framework that orchestrates planning, execution, and verification agents to decompose, ground, and iteratively refine reasoning steps over geological and hyperspectral image sets.
Outcome: The proposed framework decomposes, ground, and iteratively refines reasoning steps over geological and hyperspectral image sets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations