Papers by Bernt Schiele
More Images, More Problems? A Controlled Analysis of VLM Failure Modes. (2026.findings-acl)
Copied to clipboard
Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez
| Challenge: | Existing evaluations of large vision language models lack a comprehensive analysis of their weaknesses and causes. |
| Approach: | They propose a new benchmark to evaluate multi-image capabilities of Large Vision Language Models. |
| Outcome: | The proposed model outperforms existing benchmarks on multi-image models. |
Visual Coherence Loss for Coherent and Visually Grounded Story Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing visual storytelling models fail to generate correct referring expressions for characters, causing 60% of the generated stories to be lacking local coherence. |
| Approach: | They propose a loss function inspired by a linguistic theory of coherence for self-supervised learning for image sequence representations and a feature matching metric to check whether the models generate referring expressions correctly for characters in input image sequences. |
| Outcome: | The proposed features and loss function are effective for generating more coherent and visually grounded stories. |
Visual Writing Prompts: Character-Grounded Story Generation with Curated Image Sequences (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing work on image-based story generation lacks coherent plots for story generation. |
| Approach: | They propose to use image sequences to generate stories from a dataset that has more coherent plots. |
| Outcome: | The proposed model produces more coherent, visually grounded and diverse stories than existing models. |
A vision-grounded dataset for predicting typical locations for verbs (L18-1)
Copied to clipboard
| Challenge: | Existing models for inferring location from text are often underestimating the probability of the most typical role fillers. |
| Approach: | They propose a dataset which contains thematic fit judgments for 2,000 verb/location pairs. |
| Outcome: | The proposed dataset can be used to evaluate text-based, vision-based or multimodal inference systems for the typicality of an event's location. |