Papers by Asli Ozyurek
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue (2025.findings-acl)
Copied to clipboard
| Challenge: | Using representational co-speech gestures, face-to-face interaction participants resolve references to objects using speech and gestures. |
| Approach: | They propose a multimodal reference resolution task centred on representational gestures . they propose 'self-supervised' pre-training approach to gesture representation learning that grounds body movements in spoken language. |
| Outcome: | The proposed approach aligns with expert annotations and has significant predictive power. |
Using Perspectival Words Is Harder Than Vocabulary Words for Humans —and Even More So for Multimodal Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of multimodal language models focus on vocabulary words with relatively stable, context-independent meanings in conversation, such as object names, colors, and verbs. |
| Approach: | They compare human and multimodal language models in their use of three word types: vocabulary, possessives, and demonstratives. |
| Outcome: | The models approach human-level performance on using vocabulary, but exhibit clear deficits with possessives and even greater difficulties with demonstratives. |
The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form–Meaning Mapping (2026.acl-long)
Copied to clipboard
| Challenge: | a visual Iconicity test is used to evaluate vision–language models based on visual form and iconicity ratings. |
| Approach: | They propose a video-based benchmark to evaluate vision–language models on three tasks . they assess 17 state-of-the-art VLMs in zero- and few-shot settings on Sign Language of the Netherlands . |
| Outcome: | The proposed benchmark evaluates 17 state-of-the-art VLMs on Sign Language of the Netherlands . they achieve moderate to strong alignment with human iconicity ratings, but fail to infer lexical meaning from visual form alone . |