Challenge: Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing.
Approach: They present CopCo, the first eye tracking corpus of its kind for the Danish language.
Outcome: The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language.

Similar Papers

GECO-MT: The Ghent Eye-tracking Corpus of Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Despite improvements in machine translation output, remarkable differences can be observed when comparing machine translations (MT) and human translations.
Approach: They describe a corpus of eye movement data collected during natural reading of a human translation and a machine translation of . they use this corpus to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
Outcome: The proposed corpus will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
Reading Does Not Equal Reading: Comparing, Simulating and Exploiting Reading Behavior across Populations (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora of eye-tracking-while-reading corporata lack diversity, limiting their ability to include primarily native speakers.
Approach: They expand the eye-tracking-while-reading dataset CopCo by incorporating a new dataset of L2 readers with diverse L1 backgrounds.
Outcome: The extended CopCo corpus comprises neurotypical L1 and L1 readers with dyslexia as well as L2 readers reading the same materials.
DIDEC: The Dutch Image Description and Eye-tracking Corpus (C18-1)

Copied to clipboard

Challenge: Using a corpus of spoken Dutch image descriptions and eye-tracking data, we can gain a deeper understanding of the image description task, especially how visual attention is correlated with the image descriptions.
Approach: They present a corpus of spoken Dutch image descriptions paired with eye-tracking data from Free viewing and Description viewing tasks to provide an initial analysis of self-corrections in image descriptions.
Outcome: The results show that the eye-tracking data for the description viewing task is more coherent than for the free-viewing task, and that variation in image descriptions is only moderately correlated across different languages.
ZuCo 2.0: A Dataset of Physiological Recordings During Natural Reading and Annotation (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset of eye-tracking and electroencephalography captures language understanding . eye movement data provides millisecond-accurate records of where humans look when reading .
Approach: They recorded and preprocessed eye-tracking and electroencephalography data during natural reading and during annotation.
Outcome: The study combines eye-tracking and electroencephalography to capture the reading process . the data can be used to evaluate state-of-the-art machine learning systems .
Eye4Ref: A Multimodal Eye Movement Dataset of Referentially Complex Situations (2020.lrec-1)

Copied to clipboard

Challenge: Eye4Ref is a rich multimodal dataset of eye-movement recordings from referentially complex situated settings.
Approach: They present a rich multimodal dataset of eye-movement recordings from situated settings . they use linguistic labels, saccadic movement parameters and symbolic knowledge representations .
Outcome: The Eye4Ref dataset is an annotated multimodal dataset from three eyetracking studies on reference resolution and disambiguation tasks in situated settings.
Eye Tracking and NLP (2025.acl-tutorials)

Copied to clipboard

Challenge: tutorial combines eye tracking during reading with NLP . outlines how eye movements in reading can be leveraged for NLP methods .
Approach: The tutorial combines eye tracking during reading with NLP . it covers eye movements in reading, integrating eye movement data in NLP models .
Outcome: The tutorial outlines how eye movements in reading can be leveraged for NLP . it provides the essential background for conducting research on joint modeling of eye movements and text.
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages.
Approach: They found that the correlation between eye movements and native language similarity may be more complex than the original study found.
Outcome: The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study.
DDisCo: A Discourse Coherence Dataset for Danish (2022.lrec-1)

Copied to clipboard

Challenge: Discourse coherence models have been developed using randomly shuffled texts instead of highly edited and coherent data.
Approach: They propose to annotate Danish Wikipedia and Reddit for discourse coherence using real-world text instead of artificially incoherent text for training and testing models.
Outcome: The proposed model performs well on annotated texts from the Danish Wikipedia and Reddit dataset.
At a Glance: The Impact of Gaze Aggregation Views on Syntactic Tagging (D19-64)

Copied to clipboard

Challenge: Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus.
Approach: They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging.
Outcome: The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information.
Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement Patterns (2022.acl-long)

Copied to clipboard

Challenge: Existing work to predict gaze patterns during naturalistic reading has not been conducted on general text characteristics.
Approach: They propose to use two eye-tracking corpora of naturalistic reading and two language models to test their performance.
Outcome: The proposed models predict eye-tracking measures during naturalistic reading and language processing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations