An Annotation Approach for Social and Referential Gaze in Dialogue (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on eye gaze information focus on social functions and how it is used in reference resolution. |
| Approach: | They propose an approach for annotating eye gaze considering its social and referential functions in multi-modal dialogue. |
| Outcome: | The proposed annotation scheme is based on eye gaze behavior cues in human-human dialogues. |
Similar Papers
Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production (2024.lrec-main)
Copied to clipboard
| Challenge: | Xu and Stone et al., 2014, show eye movements are correlated with discourse goals but the relationship between eye movements and coherence is a missing link. |
| Approach: | They propose an eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study eye gaze patterns and coherence relations in multimodal language contexts. |
| Outcome: | The proposed method combines eye-tracking and a semantic gaze visualization technique to study eye movements in multimodal language contexts. |
Modeling Referential Gaze in Task-oriented Settings of Varying Referential Complexity (2022.findings-aacl)
Copied to clipboard
| Challenge: | Referential gaze is a fundamental phenomenon for psycholinguistics and human-human communication. |
| Approach: | They propose a multimodal NLP task to predict when the gaze is referential . they train a sequential attention-based LSTM model and a transformer encoder architecture to model referential gaze and transfer gaze features to unseen situated settings . |
| Outcome: | The proposed model can be applied to situations with different referential complexities . the proposed model is based on an attention-based LSTM model and a multivariate transformer encoder architecture . |
A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)
Copied to clipboard
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson
| Challenge: | Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener. |
| Approach: | They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen. |
| Outcome: | The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object. |
Dialogue Structure Annotation for Multi-Floor Interaction (L18-1)
Copied to clipboard
David Traum, Cassidy Henry, Stephanie Lukin, Ron Artstein, Felix Gervits, Kimberly Pollard, Claire Bonial, Su Lei, Clare Voss, Matthew Marge, Cory Hayes, Susan Hill
| Challenge: | Existing annotation schemes do not address dialogue structure. |
| Approach: | They propose an annotation scheme for meso-level dialogue structure that clusters utterances from multiple participants and floors into units according to realization of an initiator's intent. |
| Outcome: | The proposed annotation scheme is used to annotate a corpus of human-robot interaction dialogues. |
Dialogue Act Annotation in a Multimodal Corpus of First Encounter Dialogues (2020.lrec-1)
Copied to clipboard
| Challenge: | a method used to annotate dialogue acts in a multimodal corpus is described . the annotations allow for analysis of how multimodal signals contribute to the structure and content of the dialogues. |
| Approach: | They propose to annotate dialogue acts in a multimodal corpus of first encounter dialogues . they focus on which dialogue acts often follow each other across speakers and which overlap gestural behaviour . |
| Outcome: | The method used to annotate dialogue acts in a multimodal corpus is described. |
Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze (2020.emnlp-main)
Copied to clipboard
| Challenge: | a long tradition of cognitive studies shows that the interplay between language and vision is complex. |
| Approach: | They propose an approach to image description generation where visual processing is modelled sequentially. |
| Outcome: | The proposed model exploits gaze-driven attention to produce better descriptions . it sheds light on human cognitive processes by comparing different ways of aligning gaze with language production. |
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)
Copied to clipboard
| Challenge: | a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information. |
| Approach: | They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants. |
| Outcome: | The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants. |
A Two-Level Interpretation of Modality in Human-Robot Dialogue (2020.coling-main)
Copied to clipboard
| Challenge: | modal expressions are used to communicate and align world knowledge, but there is no obvious manner to ground them in the shared environment. |
| Approach: | They propose a two-level annotation scheme for modality that captures both content and intent and a task-oriented, pragmatic representation that maps to our robot's capabilities. |
| Outcome: | The proposed model can be grounded and dynamically interpreted. |
What Did You Refer to? Evaluating Co-References in Dialogue (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing neural end-to-end dialogue models have limitations on exactly interpreting the linguistic structures in dialogue history context. |
| Approach: | They propose to directly measure the capability of neural end-to-end dialogue models on understanding the entity-oriented structures via question answering. |
| Outcome: | The proposed model can understand large-scale English and Chinese human human dialogues using a large-format dataset. |
Grounding Language in Multi-Perspective Referential Communication (2024.emnlp-main)
Copied to clipboard
| Challenge: | Using a dataset of 2,970 human-written referring expressions, we find that the performance of automated models in both reference generation and comprehension lags behind that of pairs of human agents. |
| Approach: | They propose a task and dataset for referring expression generation and comprehension in multi-agent embodied environments where two agents must take into account one another's visual perspective to produce and understand references to objects in a scene. |
| Outcome: | The proposed model outperforms the strongest proprietary model and improves communicative success from 58.9 to 69.3% when trained with a listener. |