Papers by Haigang Zhang
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for converting visual tokens into tokens are limited by their high volume . et al., 2023; Zheng e.t., 2023): a revolution in video understanding. |
| Approach: | They propose a language-aware dynamic token compression system that converts video clips into soft caption tokens as visual representations. |
| Outcome: | The proposed method reduces FLOPs by 49% while maintaining competitive performance. |