Challenge: Existing video large language models (LMMs) employ an impedance of thousands of frames to understand long videos.
Approach: They propose a plug-and-play module integrated with VideoLLMs to facilitate efficient lengthy video perception.
Outcome: The proposed module boosts the performance of open-source VideoLLMs and proprietary assistants on long-form video benchmarks.

Similar Papers

ViLL-E: Video LLM Embeddings for Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Video Large Language Models excel at video understanding tasks where outputs are textual . however, they underperform specialized embedding-based models in Retrieval tasks .
Approach: They propose a video-LLM-based model with an embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones.
Outcome: The proposed model outperforms specialized embedding-based models in video understanding tasks while remaining competitive on VideoQA tasks.
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA (2026.eacl-long)

Copied to clipboard

Challenge: Recent studies leverage large language models (LLMs) in LVQA benchmarks, achieving exceptional performance while relying on vision language models to convert all visual content into natural language.
Approach: They propose a modular and training-free framework that leverages large language models to generate a small subset of informative frames tailored to each question.
Outcome: The proposed framework achieves state-of-the-art performance among similar models across four benchmark LVQA datasets: EgoSchema, NExT-QA, IntentQA, VideoMME.
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that Video Large Language Models (Vide-oLLMs) are efficient at video understanding but lack the quadratic complexity of visual tokens.
Approach: They propose a plug-and-play inference acceleration framework for VideoLLM token compression that quantifies each frame’s uniqueness and adaptively adjusts compression intensity across frames.
Outcome: Extensive experiments on video large language models and benchmarks show that the proposed framework can preserve essential information while reducing redundancy in video sequences.
VideoPro: Adaptive Program Reasoning for Long Video Understanding (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for understanding long videos are limited due to the sparsity of visual evidence relevant to a given query.
Approach: They propose a framework that enables VideoLLMs to reason over long videos and refine their predictions through executable programs.
Outcome: The proposed framework outperforms existing methods across long-video understanding benchmarks.
ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to adapt image-focused models for video understanding have not been successful in analyzing long video sequences.
Approach: They propose a video instruction dataset that outperforms existing video instruction data for fine-tuning MLLMs by incrementally increasing input context length.
Outcome: The proposed model outperforms existing models on video benchmarks and outperformed proprietary models on VideoMME even with a compact 7B model.
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Despite advances in reinforcement learning, data collection and fine-tuning remain costly and hard to scale.
Approach: They propose a video-adaptive test-time scaling strategy that combines RL with a supervised fine-tuning strategy to improve video reasoning capability.
Outcome: The proposed method surpasses existing models by 2.4% in accuracy using only 3.6% training samples.
Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video Understanding (2025.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models (LLMs) have shown impressive performance on the video understanding task, but extremely long videos still pose significant challenges to their context length, memory consumption, and computational complexity.
Approach: They propose a vision-language model named Sophia for long video understanding which can efficiently handle hour-scale long videos.
Outcome: The proposed model exhibits competitive performance compared to existing video understanding baselines across various benchmarks for long video understanding with reduced time and memory consumption.
InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks fail to test the full range of cognitive skills needed to process long-form videos .
Approach: They propose a benchmark to evaluate models' ability to process long-form videos rigorously.
Outcome: The benchmark measures the cognitive skills of models in understanding long-form videos . it offers the largest set of question-answer pairs for long video comprehension .
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models (2024.acl-long)

Copied to clipboard

Challenge: a surge of deep learning applications for video understanding have led to major advancements in video-related tasks.
Approach: They propose a multimodal video-based conversation model that merges a video-adapted visual encoder with an LLM and a dataset that is easily scalable and robust to label noise.
Outcome: The proposed model can understand and generate detailed conversations about videos.
VideoMind: Thinking in Steps for Long Video Understanding (2026.eacl-industry)

Copied to clipboard

Challenge: Multimodal Large Language Models struggle with Long Video Understanding due to their limited context window and the distributed nature of salient information across many redundant frames.
Approach: They propose a training framework that mimics a human reasoning process to train Long Video Understanding models.
Outcome: The proposed framework achieves 77.6% performance on Video MME, LongVideo, and MLVU benchmarks while yielding 5% improvement on Llama 4 Scout.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations