Papers by Samyak Jha
CAPA: Contribution-Aware Pruning and FFN Approximation for Efficient Large Vision-Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Efficient inference in Large Vision Language Models is constrained by the high cost of processing thousands of visual tokens. |
| Approach: | They propose a framework that prunes visual tokens using attention contribution at critical functional transitions and reduces computations using efficient linear approximations. |
| Outcome: | The proposed framework achieves competent efficiency–performance trade-offs with improved robustness. |