| Challenge: | Recent work uses gaze data at the type level or at the token level and mostly from a single eye-tracking corpus. |
| Approach: | They propose to use gaze data to capture central tendency or variability of gaze data and to integrate binary phrase chunking and part-of-speech tagging. |
| Outcome: | The proposed approaches capture the central tendency or variability of gaze data better than proposed local views which retain individual participant information. |
Similar Papers
Towards Making a Dependency Parser See (D19-1)
Copied to clipboard
| Challenge: | Eye trackers and gaze features collected from them have been recently applied to natural language processing (NLP) tasks such as part-of-speech tagging. |
| Approach: | They propose to leverage eye-tracking data in an RNN dependency parser when no aggregated or token-level gaze features are used at inference time. |
| Outcome: | The proposed model can be used to improve performance on non-gazed treebanks. |
Entity Recognition at First Sight: Improving NER with Eye Movement Information (N19-1)
Copied to clipboard
| Challenge: | Previous studies have shown eye-tracking data can be used to improve natural language processing models. |
| Approach: | They leverage eye movement features from three corpora with recorded gaze information to augment a neural model for named entity recognition with gaze embeddings. |
| Outcome: | The proposed model outperforms baseline models on both individual datasets and in cross-domain settings. |
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)
Copied to clipboard
| Challenge: | integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space. |
| Approach: | They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models. |
| Outcome: | The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space. |
Classifying Referential and Non-referential It Using Gaze (D18-1)
Copied to clipboard
| Challenge: | a particular problem for anaphora resolution systems is the pronoun it, which can be used both referentially and non-referentially. |
| Approach: | They use eye-tracking data to learn how humans perform disambiguation and use it to improve automatic classification. |
| Outcome: | The proposed system outperforms a baseline and outperformed linguistic-based approaches. |
Reading Does Not Equal Reading: Comparing, Simulating and Exploiting Reading Behavior across Populations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing corpora of eye-tracking-while-reading corporata lack diversity, limiting their ability to include primarily native speakers. |
| Approach: | They expand the eye-tracking-while-reading dataset CopCo by incorporating a new dataset of L2 readers with diverse L1 backgrounds. |
| Outcome: | The extended CopCo corpus comprises neurotypical L1 and L1 readers with dyslexia as well as L2 readers reading the same materials. |
Eye Tracking and NLP (2025.acl-tutorials)
Copied to clipboard
| Challenge: | tutorial combines eye tracking during reading with NLP . outlines how eye movements in reading can be leveraged for NLP methods . |
| Approach: | The tutorial combines eye tracking during reading with NLP . it covers eye movements in reading, integrating eye movement data in NLP models . |
| Outcome: | The tutorial outlines how eye movements in reading can be leveraged for NLP . it provides the essential background for conducting research on joint modeling of eye movements and text. |
Mind Your Special Tokens! On the Importance of Dedicated Sequence-End Tokens in Vision-Language Embedding Models (2026.eacl-short)
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) are highly sensitive to end-of-input artifacts in fine-tuning and inference data, e.g., whether input sequences end with punctuation or newline characters. |
| Approach: | They propose to convert generative LVLMs into vision-language encoders via contrastive learning objectives and use supervised contrastive objectives to train them. |
| Outcome: | The proposed approach improves visual and text representations and improves retrieval and (semantic) similarity tasks. |
Mutual Gaze and Linguistic Repetition in a Multimodal Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | a study of linguistic repetitions and mutual understanding is conducted . we find no compelling correlation between mutual gaze and duration of the event . |
| Approach: | They investigate the correlation between mutual gaze and linguistic repetition, a form of alignment, which they take as evidence of mutual understanding. |
| Outcome: | The proposed method is based on the Multisimo corpus, a multimodal corpus which provides authentic task-based interactions among three participants. |
Measuring the Impact of (Psycho-)Linguistic and Readability Features and Their Spill Over Effects on the Prediction of Eye Movement Patterns (2022.acl-long)
Copied to clipboard
| Challenge: | Existing work to predict gaze patterns during naturalistic reading has not been conducted on general text characteristics. |
| Approach: | They propose to use two eye-tracking corpora of naturalistic reading and two language models to test their performance. |
| Outcome: | The proposed models predict eye-tracking measures during naturalistic reading and language processing. |
SNAG: Spoken Narratives and Gaze Dataset (P18-2)
Copied to clipboard
| Challenge: | Existing datasets that combine gaze and spoken descriptions of visual inputs are needed to provide insight into how humans process information and make decisions. |
| Approach: | They propose a multimodal gaze and spoken descriptions dataset that can be used to label important image regions with appropriate linguistic labels. |
| Outcome: | The proposed dataset can be used to label image regions with appropriate linguistic labels. |