DIDEC: The Dutch Image Description and Eye-tracking Corpus (C18-1)

Copied to clipboard

Challenge: Using a corpus of spoken Dutch image descriptions and eye-tracking data, we can gain a deeper understanding of the image description task, especially how visual attention is correlated with the image descriptions.
Approach: They present a corpus of spoken Dutch image descriptions paired with eye-tracking data from Free viewing and Description viewing tasks to provide an initial analysis of self-corrections in image descriptions.
Outcome: The results show that the eye-tracking data for the description viewing task is more coherent than for the free-viewing task, and that variation in image descriptions is only moderately correlated across different languages.

Similar Papers

GECO-MT: The Ghent Eye-tracking Corpus of Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Despite improvements in machine translation output, remarkable differences can be observed when comparing machine translations (MT) and human translations.
Approach: They describe a corpus of eye movement data collected during natural reading of a human translation and a machine translation of . they use this corpus to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
Outcome: The proposed corpus will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes (2024.eacl-long)

Copied to clipboard

Challenge: Existing models of visuo-linguistic variation are weak to moderately trained to capture such a variation in visual outputs.
Approach: They use a corpus of Dutch image descriptions with eye-tracking data to investigate the nature of the variation in visuo-linguistic signals.
Outcome: The proposed model lacks biases about what makes a stimulus complex for humans and what leads to variations in human outputs.
The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)

Copied to clipboard

Challenge: Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing.
Approach: They present CopCo, the first eye tracking corpus of its kind for the Danish language.
Outcome: The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language.
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: The goal of the project Multilingual Image Corpus is to provide a large image dataset with annotated objects and object descriptions in 24 languages.
Approach: They propose to provide a large image dataset with annotated objects and object descriptions in 24 languages.
Outcome: The project provides a large image dataset with annotated objects and object descriptions in 24 languages.
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)

Copied to clipboard

Challenge: a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images .
Approach: They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions .
Outcome: The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces.
An Evaluation of Image-Based Verb Prediction Models against Human Eye-Tracking Data (N18-2)

Copied to clipboard

Challenge: Recent research in language and vision has developed models for predicting and disambiguating verbs from images.
Approach: They propose a verb prediction model and visual sense disambiguation model for verbs . they ask whether the image regions a model identifies as salient correlate with human intuitions about visual verbs.
Outcome: The proposed model can predict verbs from images, but it is unclear to what extent it captures human intuitions about visual verbs.
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)

Copied to clipboard

Challenge: integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
Approach: They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models.
Outcome: The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning.
Approach: They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech.
Outcome: The proposed model trains on a manually-annotated and smaller dataset does better on the task.
Every word counts: A multilingual analysis of individual human alignment with model attention (2022.aacl-short)

Copied to clipboard

Challenge: Using eye-tracking data, fixation durations are often not considered in generalisation studies because of individual differences.
Approach: They analyse eye-tracking data from speakers of 13 different languages reading . they find significant differences between languages but also individual reading behaviour .
Outcome: The proposed model can be used to improve the generalization of ML models and allow for more personalized and fair applications.
At a Glance: The Impact of Gaze Aggregation Views on Syntactic Tagging (D19-64)

Copied to clipboard

Challenge: Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus.
Approach: They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging.
Outcome: The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations