Papers by Adrian Bulat
More Images, More Problems? A Controlled Analysis of VLM Failure Modes. (2026.findings-acl)
Copied to clipboard
Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez
| Challenge: | Existing evaluations of large vision language models lack a comprehensive analysis of their weaknesses and causes. |
| Approach: | They propose a new benchmark to evaluate multi-image capabilities of Large Vision Language Models. |
| Outcome: | The proposed model outperforms existing benchmarks on multi-image models. |
Efficient Vision-Language pre-training via domain-specific learning for human activities (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current vision-language models owe their success to large-scale pretraining on web-collected data. |
| Approach: | They propose a domain-aligned pretraining strategy that aligns the downstream tasks to the downstream domain without additional data collection. |
| Outcome: | The proposed method outperforms existing models on large-scale vision-language training datasets while preserving generalist knowledge. |
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Contrastively trained Vision-Language Models exhibit shallow language understanding, manifesting bag-of-words behaviour. |
| Approach: | They propose a vision-free, single-encoder retrieval pipeline to replace traditional text-to-image retrieval paradigm with structured image descriptions. |
| Outcome: | The proposed approach reduces the modality gap and improves compositionality and performance on short and long caption queries. |