Challenge: Eye4Ref is a rich multimodal dataset of eye-movement recordings from referentially complex situated settings.
Approach: They present a rich multimodal dataset of eye-movement recordings from situated settings . they use linguistic labels, saccadic movement parameters and symbolic knowledge representations .
Outcome: The Eye4Ref dataset is an annotated multimodal dataset from three eyetracking studies on reference resolution and disambiguation tasks in situated settings.

Similar Papers

A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)

Copied to clipboard

Challenge: Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener.
Approach: They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen.
Outcome: The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object.
The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)

Copied to clipboard

Challenge: Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing.
Approach: They present CopCo, the first eye tracking corpus of its kind for the Danish language.
Outcome: The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language.
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language.
Approach: They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Outcome: The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language.
GECO-MT: The Ghent Eye-tracking Corpus of Machine Translation (2022.lrec-1)

Copied to clipboard

Challenge: Despite improvements in machine translation output, remarkable differences can be observed when comparing machine translations (MT) and human translations.
Approach: They describe a corpus of eye movement data collected during natural reading of a human translation and a machine translation of . they use this corpus to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
Outcome: The proposed corpus will be used in future research to investigate the effect of machine translation on the reading process and the effects of various error types on reading.
SNAG: Spoken Narratives and Gaze Dataset (P18-2)

Copied to clipboard

Challenge: Existing datasets that combine gaze and spoken descriptions of visual inputs are needed to provide insight into how humans process information and make decisions.
Approach: They propose a multimodal gaze and spoken descriptions dataset that can be used to label important image regions with appropriate linguistic labels.
Outcome: The proposed dataset can be used to label image regions with appropriate linguistic labels.
Linguistic, Kinematic and Gaze Information in Task Descriptions: The LKG-Corpus (2020.lrec-1)

Copied to clipboard

Challenge: linguistic structure of utterances referring to concrete actions may reflect the structure of the sensorimotor processing underlying the same action.
Approach: They present a dataset that integrates linguistic, kinematic and gaze data with an explicit focus on relations between action and language.
Outcome: The proposed dataset integrates linguistic, kinematic and gaze data with an explicit focus on relations between action and language.
Using Eye-tracking Data to Predict the Readability of Brazilian Portuguese Sentences in Single-task, Multi-task and Sequential Transfer Learning Approaches (2020.coling-main)

Copied to clipboard

Challenge: Sentence complexity assessment is a relatively new task in Natural Language Processing.
Approach: They propose to use Brazilian Portuguese to evaluate sentences with linguistic features to improve readability.
Outcome: The proposed model reaches the state-of-the-art for Brazilian Portuguese with 97.8% accuracy with linguistic features.
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
Modeling Referential Gaze in Task-oriented Settings of Varying Referential Complexity (2022.findings-aacl)

Copied to clipboard

Challenge: Referential gaze is a fundamental phenomenon for psycholinguistics and human-human communication.
Approach: They propose a multimodal NLP task to predict when the gaze is referential . they train a sequential attention-based LSTM model and a transformer encoder architecture to model referential gaze and transfer gaze features to unseen situated settings .
Outcome: The proposed model can be applied to situations with different referential complexities . the proposed model is based on an attention-based LSTM model and a multivariate transformer encoder architecture .
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied.
Approach: They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions.
Outcome: The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations