Papers by Ronja Utescher
WikiScenes with Descriptions: Aligning Paragraphs and Sentences with Images in Wikipedia Articles (2024.starsem-1)
Copied to clipboard
| Challenge: | Existing work on processing image-text alignment in multimodal documents has been unsupervised, facing the challenge of missing evaluation and training data. |
| Approach: | They propose to provide one of the first datasets that provides ground-truth annotations of image-text alignments in multi-paragraph multi-image articles. |
| Outcome: | The proposed dataset can be used to study phenomena of visual language grounding in longer documents and assess retrieval capabilities of language models trained on captioning data. |
The Illusion of Competence: Evaluating the Effect of Explanations on Users’ Mental Models of Visual Question Answering Systems (2024.emnlp-main)
Copied to clipboard
Judith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari, Heiko Wersing, Hendrik Buschmeier, Sina Zarrieß
| Challenge: | Using visual inputs, we hypothesize that explanations will make limited AI capabilities more transparent to users, but our results show that explanation increases users’ perceptions of the system’s competence regardless of its actual performance. |
| Approach: | They employ a visual question answer and explanation task where participants control the AI system’s limitations by manipulating visual inputs. |
| Outcome: | The proposed explanations do not increase users’ perceptions of the system’s competence regardless of its actual performance. |