Challenge: Despite improvements in machine translation output, remarkable differences can be observed when comparing machine translations (MT) and human translations.
Approach: They describe a corpus of eye movement data collected during natural reading of a human translation and a machine translation of . they use this corpus to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
Outcome: The proposed corpus will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading.

Similar Papers

The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)

Copied to clipboard

Challenge: Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing.
Approach: They present CopCo, the first eye tracking corpus of its kind for the Danish language.
Outcome: The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language.
DIDEC: The Dutch Image Description and Eye-tracking Corpus (C18-1)

Copied to clipboard

Challenge: Using a corpus of spoken Dutch image descriptions and eye-tracking data, we can gain a deeper understanding of the image description task, especially how visual attention is correlated with the image descriptions.
Approach: They present a corpus of spoken Dutch image descriptions paired with eye-tracking data from Free viewing and Description viewing tasks to provide an initial analysis of self-corrections in image descriptions.
Outcome: The results show that the eye-tracking data for the description viewing task is more coherent than for the free-viewing task, and that variation in image descriptions is only moderately correlated across different languages.
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)

Copied to clipboard

Challenge: integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
Approach: They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models.
Outcome: The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
On Context Span Needed for Machine Translation Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: a number of common patterns can be observed for context-aware MT evaluation, authors say . document-level evaluations have largely been performed at the sentence level . the definition of what constitutes a "document level" evaluation is still unclear .
Approach: They propose to use a series of surveys to identify the necessary context span . they find common patterns that can be used to draw general guidelines .
Outcome: The proposed evaluations of machine translation systems show that some issues and spans depend on domain and target language.
Selecting Machine-Translated Data for Quick Bootstrapping of a Natural Language Understanding System (N18-3)

Copied to clipboard

Challenge: In recent years, there has been growing interest in voice-controlled devices, such as Amazon Alexa or Google home.
Approach: They investigate the use of Machine Translation to bootstrap a natural language understanding system for a new language for the use case of a large-scale voice-controlled device.
Outcome: The proposed method reduces the time and cost of getting annotated corpus for a new language while still providing a large enough coverage of user requests.
DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of human translations contains both professional and student translations of news and reviews texts.
Approach: They propose to use the data to compare human and professional translations of news and reviews in a new corpus which contains both professional and student translations.
Outcome: The proposed corpus contains professional and student translations of news and reviews and a subcorpus containing reviews into Finnish.
ZuCo 2.0: A Dataset of Physiological Recordings During Natural Reading and Annotation (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset of eye-tracking and electroencephalography captures language understanding . eye movement data provides millisecond-accurate records of where humans look when reading .
Approach: They recorded and preprocessed eye-tracking and electroencephalography data during natural reading and during annotation.
Outcome: The study combines eye-tracking and electroencephalography to capture the reading process . the data can be used to evaluate state-of-the-art machine learning systems .
Eye Tracking and NLP (2025.acl-tutorials)

Copied to clipboard

Challenge: tutorial combines eye tracking during reading with NLP . outlines how eye movements in reading can be leveraged for NLP methods .
Approach: The tutorial combines eye tracking during reading with NLP . it covers eye movements in reading, integrating eye movement data in NLP models .
Outcome: The tutorial outlines how eye movements in reading can be leveraged for NLP . it provides the essential background for conducting research on joint modeling of eye movements and text.
Eye Movement Features Can Predict Human Preferences on Machine-Generated Texts (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies on eye movement in text quality assessment are limited . eye-movement features are important predictors of human judgments of text quality, but are costly and inconsistent.
Approach: They propose to capture eye-movement features during screen reading of LLM-generated text using a dataset that includes eye-motion recordings, reading-time measurements, and post-reading evaluations.
Outcome: The proposed dataset shows that eye-movement features can significantly improve models over other probabilistic metrics, including negative log-likelihood (NLL).
Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World Literature (2022.emnlp-main)

Copied to clipboard

Challenge: Literary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators . a dataset of non-English language novels is used to study literary MT .
Approach: They use a dataset of non-English language novels aligned to human and automatic English translations to study literary MT.
Outcome: The proposed model prefers human translations over machine translations at a rate of 84% . state-of-the-art MT metrics do not correlate with preferences, the study finds .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations