Challenge: Existing corpora of eye-tracking-while-reading corporata lack diversity, limiting their ability to include primarily native speakers.
Approach: They expand the eye-tracking-while-reading dataset CopCo by incorporating a new dataset of L2 readers with diverse L1 backgrounds.
Outcome: The extended CopCo corpus comprises neurotypical L1 and L1 readers with dyslexia as well as L2 readers reading the same materials.

Similar Papers

The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)

Copied to clipboard

Challenge: Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing.
Approach: They present CopCo, the first eye tracking corpus of its kind for the Danish language.
Outcome: The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language.
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)

Copied to clipboard

Challenge: integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
Approach: They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models.
Outcome: The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
Eye Tracking and NLP (2025.acl-tutorials)

Copied to clipboard

Challenge: tutorial combines eye tracking during reading with NLP . outlines how eye movements in reading can be leveraged for NLP methods .
Approach: The tutorial combines eye tracking during reading with NLP . it covers eye movements in reading, integrating eye movement data in NLP models .
Outcome: The tutorial outlines how eye movements in reading can be leveraged for NLP . it provides the essential background for conducting research on joint modeling of eye movements and text.
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages.
Approach: They found that the correlation between eye movements and native language similarity may be more complex than the original study found.
Outcome: The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study.
Entity Recognition at First Sight: Improving NER with Eye Movement Information (N19-1)

Copied to clipboard

Challenge: Previous studies have shown eye-tracking data can be used to improve natural language processing models.
Approach: They leverage eye movement features from three corpora with recorded gaze information to augment a neural model for named entity recognition with gaze embeddings.
Outcome: The proposed model outperforms baseline models on both individual datasets and in cross-domain settings.
Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement Patterns (2022.acl-long)

Copied to clipboard

Challenge: Existing work to predict gaze patterns during naturalistic reading has not been conducted on general text characteristics.
Approach: They propose to use two eye-tracking corpora of naturalistic reading and two language models to test their performance.
Outcome: The proposed models predict eye-tracking measures during naturalistic reading and language processing.
InteRead: An Eye Tracking Dataset of Interrupted Reading (2024.lrec-main)

Copied to clipboard

Challenge: Eye movements during reading can provide insights into cognitive processes and language comprehension, but the scarcity of reading data with interruptions hampers advances in the development of intelligent learning technologies.
Approach: They propose a dataset of eye movements during reading that includes eye movements and word frequency effects.
Outcome: The proposed dataset shows that interruptions, word length and word frequency effects significantly impact eye movements during reading.
At a Glance: The Impact of Gaze Aggregation Views on Syntactic Tagging (D19-64)

Copied to clipboard

Challenge: Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus.
Approach: They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging.
Outcome: The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information.
Multilingual Language Models Predict Human Reading Behavior (2021.naacl-main)

Copied to clipboard

Challenge: Recent studies show that cognitively motivated "attention" mechanism in neural models is not a good indicator for relative importance.
Approach: They compare the performance of language-specific and multilingual pretrained transformer models to predict reading time measures reflecting natural human sentence processing.
Outcome: The proposed models predict reading time measures on Dutch, English, German, and Russian texts.
Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading Times (2023.acl-short)

Copied to clipboard

Challenge: Neural language models provide conditional probability distributions over the lexicon that are predictive of human processing times.
Approach: They propose to use a transformer-based model to generate probabilistic estimates that are less predictive of early eye-tracking measurements reflecting lexical access and early semantic integration.
Outcome: The proposed models show that larger models capture late eye-tracking measurements that reflect the full integration of a word into the current language context.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations