Papers by Gedas Bertasius

3 papers
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Despite advances in reinforcement learning, data collection and fine-tuning remain costly and hard to scale.
Approach: They propose a video-adaptive test-time scaling strategy that combines RL with a supervised fine-tuning strategy to improve video reasoning capability.
Outcome: The proposed method surpasses existing models by 2.4% in accuracy using only 3.6% training samples.
A Simple LLM Framework for Long-Range Video Question-Answering (2024.emnlp-main)

Copied to clipboard

Challenge: a recent study has shown that short video understanding is not trivial due to the need for long-range temporal reasoning capabilities.
Approach: They propose a language-based short- and long-range question-answering framework LLoVi . they propose 'multi-round summarization prompt' that asks the LLM to summarize the captions .
Outcome: The proposed framework outperforms the state-of-the-art on the EgoSchema dataset and to grounded VideoQA.
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences (2024.acl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks, but their ability to extrapolate from image sequences has been less investigated.
Approach: They propose a new benchmark to assess MLLMs’ sequential image reasoning abilities.
Outcome: The proposed benchmark features 4,761 diverse image sequences with varying lengths.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations