Papers by Ilker Yildirim
When are Lemons Purple? The Concept Association Bias of Vision-Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-to-text retrieval tasks. |
| Approach: | They propose to use "question text" as input for the text encoder of CLIP to make the prediction harder than it should be. |
| Outcome: | The proposed model treats input as a bag of concepts and attempts to fill in the other missing concept crossmodally, leading to an unexpected zero-shot prediction. |