Papers with SigLIP

3 papers
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Recent work reveals that vision and language models struggle to comprehend fine grained distinctions in images.
Approach: They propose a dataset to assess multimodal models' ability to match objects with their colors.
Outcome: The proposed model performs well in visual questionanswering, text-to-image generation and word-order understanding tasks.
Nearest Neighbor Normalization Improves Multimodal Retrieval (2024.emnlp-main)

Copied to clipboard

Challenge: Recent training-free methods suggest that accuracy can be improved without fine-tuning.
Approach: They propose a method for correcting errors in trained contrastive image-text retrieval models with no additional training, called Nearest Neighbor Normalization.
Outcome: The proposed method improves retrieval metrics for all contrastive models and datasets and does not require training on the reference database.
Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Dataset (2025.emnlp-main)

Copied to clipboard

Challenge: Current evaluations for Vision-language Models remain heavily anchored to ImageNet .
Approach: They propose a large-scale semantically-annotated multimodal resource that extends the range of visual concepts, including diverse abstract categories.
Outcome: The proposed model expands the range of visual concepts, including diverse abstract categories.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations