Papers by Rishika Bhagwatkar
Improving Adversarial Robustness in Vision-Language Models with Architecture and Prompt Design (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have seen a significant increase in research interest and real-world applications, including healthcare, autonomous systems, and security. |
| Approach: | They propose novel approaches to enhance model robustness through prompt engineering by suggesting adversarial perturbations or rephrasing questions. |
| Outcome: | The proposed approaches improve model robustness against strong image-based attacks such as Auto-PGD. |
CAVE : Detecting and Explaining Commonsense Anomalies in Visual Environments (2025.emnlp-main)
Copied to clipboard
Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges, Irina Rish, Antoine Bosselut
| Challenge: | a new benchmark for computer vision fails to capture richness and unpredictability of real-world anomalies . state-of-the-art VLMs struggle with visual anomaly perception and commonsense reasoning . elucidating the nature of anomalies is a fundamental human trait . |
| Approach: | They propose a benchmark for visual anomalies that includes annotations for visual grounding and categorizing anomalies based on their visual manifestations, their complexity, severity, and commonness. |
| Outcome: | The proposed benchmark improves on existing vision models by incorporating visual annotations. |