Papers by Moshiur Farazi
CAPSTONE: Composable Attribute‐Prompted Scene Translation for Zero‐Shot Vision–Language Reasoning (2025.emnlp-industry)
Copied to clipboard
| Challenge: | CAPSTONE transforms visual inputs into structured text prompts that can be interpreted by a frozen Large Language Model (LLM). |
| Approach: | They propose a plug-and-play framework that transforms off-the-shelf vision models into structured text prompts that can be interpreted by a frozen Large Language Model (LLM). |
| Outcome: | The proposed framework outperforms fully trained VLMs on the POPE dataset while the 4B model achieves competitive results. |