Challenge: Existing methods for video-text retrieval capture fine-grained semantic concepts . however, they lack the ability to capture finer-grain concepts such as objects and actions.
Approach: They propose a dual-encoder architecture for fast video-text retrieval that learns lexicon representations to capture fine-grained semantics.
Outcome: The proposed framework outperforms existing methods with 4.8% and 8.2% improvement on MSR-VTT and DiDeMo respectively.

Similar Papers

ViLL-E: Video LLM Embeddings for Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Video Large Language Models excel at video understanding tasks where outputs are textual . however, they underperform specialized embedding-based models in Retrieval tasks .
Approach: They propose a video-LLM-based model with an embedding generation mechanism that allows the model to "think longer" for complex videos and stop early for easy ones.
Outcome: The proposed model outperforms specialized embedding-based models in video understanding tasks while remaining competitive on VideoQA tasks.
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Video Large Language Models (VLMs) have been praised for their performance in coarse-grained video understanding but still face ineffective temporal grounding and inadequate timestamp representations.
Approach: They propose a novel Video-LLM that senses and reasoned over specific video moments with fine-grained temporal precision.
Outcome: The proposed model surpasses existing models in fine-grained video understanding tasks and exhibits strong potential as a general video understanding assistant.
Lexicon-Enhanced Self-Supervised Training for Multilingual Dense Retrieval (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent multilingual pre-trained models perform poorly on multilingual retrieval tasks due to lack of multilingual training data.
Approach: They propose to mine and generate self-supervised training data based on large-scale unlabeled corpus and introduce query generator to generate more queries in target languages for unlabed passages.
Outcome: The proposed method performs better than baselines on a Mr. TYDI dataset and an industrial dataset from a commercial search engine.
MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for video retrieval rely on embedding-based full-corpus scanning, but there is a bottleneck in semantic asymmetry and computational redundancy.
Approach: They propose a multi-agent framework that rethinks retrieval as cooperative reasoning . they parse raw videos into a structured semantic library, enabling explicit attribute-level indexing .
Outcome: The proposed framework bridges the granularity mismatch gap by parsing raw videos into a structured semantic library . it employs a Logic-aware Debate mechanism with a strict veto protocol . the proposed framework achieves competitive performance without task-specific fine-tuning .
Field Embedding: A Unified Grain-Based Framework for Word Representation (2021.naacl-main)

Copied to clipboard

Challenge: Current methods focus on learning word embeddings while linguistic information is discarded after the learning.
Approach: They propose a framework field embedding to jointly learn word and grain embedds by incorporating morphological, phonetic, and syntactical linguistic fields.
Outcome: The proposed framework integrates morphological, phonetic, and syntactical linguistic fields to learn word embeddings and grain embedds.
DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in video-text retrieval models have limited training data annotations.
Approach: They propose a Video-Text Retrieval Paradigm with Relevance-based Augmentation which enhances video and text data using large foundation models to learn more generalized features.
Outcome: The proposed method improves video-text retrieval performance over existing methods.
Video-Text Retrieval by Supervised Sparse Multi-Grained Learning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in video-text retrieval have led to improved representation learning methods.
Approach: They propose a multi-grained sparse learning framework to learn an aligned sparsen space shared between video and text for video-text retrieval.
Outcome: The proposed framework is superior to existing methods on video-text retrieval benchmarks.
Contrastive Video-Language Learning with Fine-grained Frame Sampling (2022.aacl-main)

Copied to clipboard

Challenge: despite recent progress in video and language representation learning, the weak or sparse correspondence between the two modalities remains a bottleneck.
Approach: They propose a fine-grained contrastive objective for video frame sampling to improve cross-modal correspondence.
Outcome: The proposed approach achieves state-of-the-art performance on YouCookII with long videos.
LexFit: Lexical Fine-Tuning of Pretrained Language Models (2021.acl-long)

Copied to clipboard

Challenge: Transformer-based language models implicitly store a wealth of lexical semantic knowledge, but it is non-trivial to extract that knowledge effectively from their parameters.
Approach: They propose to expose and enrich lexical knowledge from transformer-based language models to serve as effective decontextualized word encoders even when fed input words "in isolation"
Outcome: The proposed model outperforms standard static WEs and vanilla LMs in lexical tasks over four established tasks in 8 languages.
Cross-Modal Discrete Representation Learning (2022.acl-long)

Copied to clipboard

Challenge: a new framework for learning representations from multimodal data is proposed . the proposed framework uses discretized embedding vectors to capture finer levels of granularity .
Approach: They propose a self-supervised representation learning framework that captures finer levels of granularity across different modalities.
Outcome: The proposed representation can capture finer levels of granularity across different modalities . it can be used on cross-modal retrieval tasks without direct supervision .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations