Papers with SigLIP
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent work reveals that vision and language models struggle to comprehend fine grained distinctions in images. |
| Approach: | They propose a dataset to assess multimodal models' ability to match objects with their colors. |
| Outcome: | The proposed model performs well in visual questionanswering, text-to-image generation and word-order understanding tasks. |
Nearest Neighbor Normalization Improves Multimodal Retrieval (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent training-free methods suggest that accuracy can be improved without fine-tuning. |
| Approach: | They propose a method for correcting errors in trained contrastive image-text retrieval models with no additional training, called Nearest Neighbor Normalization. |
| Outcome: | The proposed method improves retrieval metrics for all contrastive models and datasets and does not require training on the reference database. |
Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Dataset (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current evaluations for Vision-language Models remain heavily anchored to ImageNet . |
| Approach: | They propose a large-scale semantically-annotated multimodal resource that extends the range of visual concepts, including diverse abstract categories. |
| Outcome: | The proposed model expands the range of visual concepts, including diverse abstract categories. |