Papers by Adrian Bulat

3 papers
More Images, More Problems? A Controlled Analysis of VLM Failure Modes. (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of large vision language models lack a comprehensive analysis of their weaknesses and causes.
Approach: They propose a new benchmark to evaluate multi-image capabilities of Large Vision Language Models.
Outcome: The proposed model outperforms existing benchmarks on multi-image models.
Efficient Vision-Language pre-training via domain-specific learning for human activities (2024.emnlp-main)

Copied to clipboard

Challenge: Current vision-language models owe their success to large-scale pretraining on web-collected data.
Approach: They propose a domain-aligned pretraining strategy that aligns the downstream tasks to the downstream domain without additional data collection.
Outcome: The proposed method outperforms existing models on large-scale vision-language training datasets while preserving generalist knowledge.
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions (2025.emnlp-main)

Copied to clipboard

Challenge: Contrastively trained Vision-Language Models exhibit shallow language understanding, manifesting bag-of-words behaviour.
Approach: They propose a vision-free, single-encoder retrieval pipeline to replace traditional text-to-image retrieval paradigm with structured image descriptions.
Outcome: The proposed approach reduces the modality gap and improves compositionality and performance on short and long caption queries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations