Papers by Nimrod Shabtay
CARES: Context-Aware Resolution Selector for VLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Large vision–language models process images at native or high resolution to remain effective across tasks. |
| Approach: | They propose a lightweight preprocessing module that predicts the minimum sufficient input resolution for large vision–language models. |
| Outcome: | CARES predicts when a pre-trained VLM's response converges to its peak ability to answer correctly, reducing compute by up to 80%. |