Papers by Zahidul Islam
AdaptMerge: Inference Time Adaptive Visual and Language-Guided Token Merging for Efficient Large Multimodal Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing token reduction methods ignore image complexity and vision-language interactions, ignoring image complexity. |
| Approach: | They propose a training-free, inference-time token merging strategy that adaptively reduces visual tokens by leveraging feature diversity and language-guided relevance. |
| Outcome: | The proposed approach outperforms state-of-the-art token reduction methods on Google’s Gemma 3 models while achieving reduced computational costs and improved performance. |