Papers by Kunzhe Huang
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that CLIP models with short summary texts cannot process extensive textual descriptions due to its text encoder's reliance on positional embeddings with length 77. |
| Approach: | They propose a Contrastive Language-Image Pre-training (CLIP) model which aims to unleash the long-description understanding capability of video CLIP models. |
| Outcome: | The proposed model can learn the distribution of feature space while expanding the long description capability. |