Papers by Andreas Bulling
InteRead: An Eye Tracking Dataset of Interrupted Reading (2024.lrec-main)
Copied to clipboard
| Challenge: | Eye movements during reading can provide insights into cognitive processes and language comprehension, but the scarcity of reading data with interruptions hampers advances in the development of intelligent learning technologies. |
| Approach: | They propose a dataset of eye movements during reading that includes eye movements and word frequency effects. |
| Outcome: | The proposed dataset shows that interruptions, word length and word frequency effects significantly impact eye movements during reading. |
ProToM: Promoting Prosocial Behaviour via Theory of Mind-Informed Feedback (2026.findings-acl)
Copied to clipboard
| Challenge: | ProToM provides targeted, context-sensitive feedback to individual agents, achieving a higher success rate, shorter task completion times, and is consistently preferred by human users. |
| Approach: | They propose a Theory of Mind-informed facilitator that provides targeted, context-sensitive feedback to individual agents. |
| Outcome: | The proposed system provides targeted, context-sensitive feedback to promote prosocial behaviour, even when not directly aligned with one’s own goals. |
Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan Acquisition (2024.acl-long)
Copied to clipboard
| Challenge: | Recent work on dialogue-based collaborative plan acquisition (CPA) suggests Theory of Mind (ToM) modelling can improve missing knowledge prediction in settings with asymmetric skill-sets and knowledge. |
| Approach: | They propose to use task-specific constraints to represent plans as graphs and exploit task-related constraints to improve missing knowledge prediction in CPA. |
| Outcome: | The proposed model improves missing knowledge prediction in contexts with asymmetric skill-sets and knowledge, but the improvements diminish . the proposed model is compared with baseline models and found to be more effective than existing models. |
ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing Theory of Mind (ToM) benchmarks focus on text-only or dyadic interactions, but to address this gap, we propose ToM-SSI: a new benchmark specifically designed to test ToM capabilities in environments rich with social interactions and spatial dynamics. |
| Approach: | They propose to use the Sally-Anne test to test ToM capabilities in environments rich in social interactions and spatial dynamics. |
| Outcome: | The proposed model captures a wider range of social cognition than existing models and demonstrates that existing models are still limited in these new tasks. |
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models, little is known about how LMs internally represent mental states of self and others. |
| Approach: | They propose to investigate how LMs internally represent mental states of self and others . |
| Outcome: | The proposed model size and finetuning significantly improve LMs’ internal representations of others’ beliefs, which are structured - not mere by-products of spurious correlations - yet brittle to prompt variations. |
Neuro-Symbolic Visual Dialog (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for visual dialog require large amounts of training data, which is prohibitive for most settings. |
| Approach: | They propose a method that integrates deep learning and symbolic program execution for multi-round visual reasoning. |
| Outcome: | The proposed model outperforms existing methods on long-distance co-reference resolution and vanishing question-answering performance. |
A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks (L18-1)
Copied to clipboard
| Challenge: | Phonetic segmentation is the process of splitting speech into distinct phonetic units . methods for automatic segmentation are not always accurate enough . |
| Approach: | They propose to model phonetic segmentation as close as possible to manual segmentation by recording experts performing a segmentation task. |
| Outcome: | This corpus captures human segmentation behavior by recording experts performing a segmentation task. |
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing video dialog models struggle with questions requiring both spatial and temporal localization within videos, long-term temporal reasoning, and accurate object tracking across multiple dialog turns. |
| Approach: | They propose a multi-modal attention-based model for video dialog operating over a dialog state tracker. |
| Outcome: | The proposed model can learn multi-modal dialog state representations of the most relevant objects and rounds. |