A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)
Copied to clipboard
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson
| Challenge: | Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener. |
| Approach: | They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen. |
| Outcome: | The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object. |
Similar Papers
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups. |
| Approach: | They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks. |
| Outcome: | The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants. |
The AICO Multimodal Corpus – Data Collection and Preliminary Analyses (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on human multimodal behaviour in interactions with a human or a robot partner are limited. |
| Approach: | They describe the first explorative research on the AICO Multimodal Corpus, which contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions. |
| Outcome: | The AICO Multimodal Corpus contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions. |
Mutual Gaze and Linguistic Repetition in a Multimodal Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | a study of linguistic repetitions and mutual understanding is conducted . we find no compelling correlation between mutual gaze and duration of the event . |
| Approach: | They investigate the correlation between mutual gaze and linguistic repetition, a form of alignment, which they take as evidence of mutual understanding. |
| Outcome: | The proposed method is based on the Multisimo corpus, a multimodal corpus which provides authentic task-based interactions among three participants. |
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)
Copied to clipboard
| Challenge: | a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information. |
| Approach: | They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants. |
| Outcome: | The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants. |
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied. |
| Approach: | They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions. |
| Outcome: | The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions. |
RoomReader: A Multimodal Corpus of Online Multiparty Conversational Interactions (2022.lrec-1)
Copied to clipboard
Justine Reverdy, Sam O’Connor Russell, Louise Duquenne, Diego Garaialde, Benjamin R. Cowan, Naomi Harte
| Challenge: | The corpus of multimodal, multiparty conversational interactions explored in RoomReader can be used to study a wide range of phenomena in online multimodal interaction. |
| Approach: | They propose to use RoomReader to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments. |
| Outcome: | The corpus was developed within the wider RoomReader Project to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments. |
MULTICOLLAB: A Multimodal Corpus of Dialogues for Analyzing Collaboration and Frustration in Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to study complex emotions when a speaker collaborates with a partner are limited. |
| Approach: | They propose to fuse a multimodal dialogue resource with transcribed speech and eye gaze data to create a highly multimodal corpus. |
| Outcome: | The proposed model improves classification accuracy by 21% over baseline using sensor and speech data in 4.5 seconds. |
Eye4Ref: A Multimodal Eye Movement Dataset of Referentially Complex Situations (2020.lrec-1)
Copied to clipboard
| Challenge: | Eye4Ref is a rich multimodal dataset of eye-movement recordings from referentially complex situated settings. |
| Approach: | They present a rich multimodal dataset of eye-movement recordings from situated settings . they use linguistic labels, saccadic movement parameters and symbolic knowledge representations . |
| Outcome: | The Eye4Ref dataset is an annotated multimodal dataset from three eyetracking studies on reference resolution and disambiguation tasks in situated settings. |
Action Verb Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of 390 simple actions is based on multimodal data of 12 humans . the dataset is annotated with orthographic transcriptions of utterances and part-of-speech tags . |
| Approach: | They present a multimodal corpus of 12 humans performing 390 simple actions . they also propose an algorithm for segmenting words into utterances and aligning visual information and speech . |
| Outcome: | The presented dataset includes 390 simple actions performed by 12 humans . it includes transcriptions of utterances, part-of-speech tags, lemmata, and hand touches . |
SNAG: Spoken Narratives and Gaze Dataset (P18-2)
Copied to clipboard
| Challenge: | Existing datasets that combine gaze and spoken descriptions of visual inputs are needed to provide insight into how humans process information and make decisions. |
| Approach: | They propose a multimodal gaze and spoken descriptions dataset that can be used to label important image regions with appropriate linguistic labels. |
| Outcome: | The proposed dataset can be used to label image regions with appropriate linguistic labels. |