Papers by Amir Atapour-Abarghouei
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion (2025.findings-emnlp)
Copied to clipboard
| Challenge: | proposed lightweight MLLM framework for end-to-end visual question answering . proposed framework uses BreezeCLIP, a vision-language encoder optimised for efficient multimodal understanding . |
| Approach: | proposed lightweight MLLM framework is based on BreezeCLIP, a vision-language encoder . it offers a promising path toward deployable ML models under practical hardware constraints. |
| Outcome: | The proposed model significantly reduces computational cost while achieving performance comparable to standard-size MLLMs. |