Papers by Simeon Junker
Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing models are unable to resolve references to abstract visual stimuli, such as color patches and color grids, but their pragmatic capabilities are still a challenge for state-of-the-art MLLMs. |
| Approach: | They investigate whether multimodal large language models are able to resolve references to abstract visual stimuli, such as color patches and color grids, in a well-known reference resolution paradigm. |
| Outcome: | The proposed model can resolve references to abstract visual stimuli in dyadic reference games. |
SceneGram: Conceptualizing and Describing Tangrams in Scene Context (2025.findings-acl)
Copied to clipboard
| Challenge: | Current systems show mixed results in reproducing human variation in object naming . figurative descriptions for abstract stimuli remain a major challenge in vision and language research . |
| Approach: | They propose to analyze human references to tangrams placed in different scene contexts . they analyze the richness and variability of conceptualizations found in human references . |
| Outcome: | The proposed model does not account for the richness and variability of human references. |
The Illusion of Competence: Evaluating the Effect of Explanations on Users’ Mental Models of Visual Question Answering Systems (2024.emnlp-main)
Copied to clipboard
Judith Sieker, Simeon Junker, Ronja Utescher, Nazia Attari, Heiko Wersing, Hendrik Buschmeier, Sina Zarrieß
| Challenge: | Using visual inputs, we hypothesize that explanations will make limited AI capabilities more transparent to users, but our results show that explanation increases users’ perceptions of the system’s competence regardless of its actual performance. |
| Approach: | They employ a visual question answer and explanation task where participants control the AI system’s limitations by manipulating visual inputs. |
| Outcome: | The proposed explanations do not increase users’ perceptions of the system’s competence regardless of its actual performance. |