Papers with SA-CLIP
SA-CLIP: Language Guided Image Spatial and Action Feature Learning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Contrastive language-image pretraining models struggle with real-world downstream tasks such as road traffic anomaly detection due to inability to effectively capture spatial and action relationships between objects within images. |
| Approach: | They compile and curate a dataset and train a Spatial and Action relationship aware CLIP model. |
| Outcome: | The proposed model performs well on the traffic anomaly detection task . |