| Challenge: | Using a corpus of spoken Dutch image descriptions and eye-tracking data, we can gain a deeper understanding of the image description task, especially how visual attention is correlated with the image descriptions. |
| Approach: | They present a corpus of spoken Dutch image descriptions paired with eye-tracking data from Free viewing and Description viewing tasks to provide an initial analysis of self-corrections in image descriptions. |
| Outcome: | The results show that the eye-tracking data for the description viewing task is more coherent than for the free-viewing task, and that variation in image descriptions is only moderately correlated across different languages. |
Similar Papers
GECO-MT: The Ghent Eye-tracking Corpus of Machine Translation (2022.lrec-1)
Copied to clipboard
| Challenge: | Despite improvements in machine translation output, remarkable differences can be observed when comparing machine translations (MT) and human translations. |
| Approach: | They describe a corpus of eye movement data collected during natural reading of a human translation and a machine translation of . they use this corpus to investigate the effect of machine translation on the reading process and the effects of various error types on reading. |
| Outcome: | The proposed corpus will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading. |
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing models of visuo-linguistic variation are weak to moderately trained to capture such a variation in visual outputs. |
| Approach: | They use a corpus of Dutch image descriptions with eye-tracking data to investigate the nature of the variation in visuo-linguistic signals. |
| Outcome: | The proposed model lacks biases about what makes a stimulus complex for humans and what leads to variations in human outputs. |
The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)
Copied to clipboard
| Challenge: | Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing. |
| Approach: | They present CopCo, the first eye tracking corpus of its kind for the Danish language. |
| Outcome: | The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language. |
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset (2022.lrec-1)
Copied to clipboard
| Challenge: | The goal of the project Multilingual Image Corpus is to provide a large image dataset with annotated objects and object descriptions in 24 languages. |
| Approach: | They propose to provide a large image dataset with annotated objects and object descriptions in 24 languages. |
| Outcome: | The project provides a large image dataset with annotated objects and object descriptions in 24 languages. |
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)
Copied to clipboard
Albert Gatt, Marc Tanti, Adrian Muscat, Patrizia Paggio, Reuben A Farrugia, Claudia Borg, Kenneth P Camilleri, Michael Rosner, Lonneke van der Plas
| Challenge: | a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images . |
| Approach: | They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions . |
| Outcome: | The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces. |
An Evaluation of Image-Based Verb Prediction Models against Human Eye-Tracking Data (N18-2)
Copied to clipboard
| Challenge: | Recent research in language and vision has developed models for predicting and disambiguating verbs from images. |
| Approach: | They propose a verb prediction model and visual sense disambiguation model for verbs . they ask whether the image regions a model identifies as salient correlate with human intuitions about visual verbs. |
| Outcome: | The proposed model can predict verbs from images, but it is unclear to what extent it captures human intuitions about visual verbs. |
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)
Copied to clipboard
| Challenge: | integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space. |
| Approach: | They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models. |
| Outcome: | The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space. |
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)
Copied to clipboard
| Challenge: | Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning. |
| Approach: | They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech. |
| Outcome: | The proposed model trains on a manually-annotated and smaller dataset does better on the task. |
Every word counts: A multilingual analysis of individual human alignment with model attention (2022.aacl-short)
Copied to clipboard
| Challenge: | Using eye-tracking data, fixation durations are often not considered in generalisation studies because of individual differences. |
| Approach: | They analyse eye-tracking data from speakers of 13 different languages reading . they find significant differences between languages but also individual reading behaviour . |
| Outcome: | The proposed model can be used to improve the generalization of ML models and allow for more personalized and fair applications. |
At a Glance: The Impact of Gaze Aggregation Views on Syntactic Tagging (D19-64)
Copied to clipboard
| Challenge: | Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus. |
| Approach: | They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging. |
| Outcome: | The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information. |