Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.

Similar Papers

Collection of Multimodal Dialog Data and Analysis of the Result of Annotation of Users’ Interest Level (L18-1)

Copied to clipboard

Challenge: a group of researchers is building a corpus for evaluating elements of multimodal dialogue systems.
Approach: They propose to build a corpus for evaluating elements of the multimodal dialogue system . they use the Wizard of Oz method to record chat dialogue data between a human and a virtual agent .
Outcome: The proposed method annotates chat dialogue data between a human and a virtual agent and measures their interest level in the data.
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.
Dialogue Act Annotation in a Multimodal Corpus of First Encounter Dialogues (2020.lrec-1)

Copied to clipboard

Challenge: a method used to annotate dialogue acts in a multimodal corpus is described . the annotations allow for analysis of how multimodal signals contribute to the structure and content of the dialogues.
Approach: They propose to annotate dialogue acts in a multimodal corpus of first encounter dialogues . they focus on which dialogue acts often follow each other across speakers and which overlap gestural behaviour .
Outcome: The method used to annotate dialogue acts in a multimodal corpus is described.
The ADELE Corpus of Dyadic Social Text Conversations:Dialog Act Annotation with ISO 24617-2 (L18-1)

Copied to clipboard

Challenge: Recent studies have focused on task-based or instrumental dialogs, but there is increasing interest in social or interactional dialogs.
Approach: They describe a corpus of 193 dyadic text dialogs based on a novel 'getting to know you' social dialog elicitation paradigm and propose additional acts to better cover greeting and leavetaking.
Outcome: The proposed actions cover greeting and leavetaking, and the proposed acts improve the interaction between the dialogs and spoken language.
Chats and Chunks: Annotation and Analysis of Multiparty Long Casual Conversations (L18-1)

Copied to clipboard

Challenge: dyadic conversations are attracting more interest with attempts to build more friendly and natural spoken dialog systems.
Approach: They describe the collection, organization, and annotation of structural chat and chunk phases in three existing corpora and analyse their preliminary results to find that chunk dominates as conversations get longer.
Outcome: The results show that chunk dominates conversations as they get longer .
MULTICOLLAB: A Multimodal Corpus of Dialogues for Analyzing Collaboration and Frustration in Language (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to study complex emotions when a speaker collaborates with a partner are limited.
Approach: They propose to fuse a multimodal dialogue resource with transcribed speech and eye gaze data to create a highly multimodal corpus.
Outcome: The proposed model improves classification accuracy by 21% over baseline using sensor and speech data in 4.5 seconds.
Multimodal Persona Based Generation of Comic Dialogs (2023.acl-long)

Copied to clipboard

Challenge: Existing models for persona based dialogue generation for comic strips encode two-party dialogues and do not account for visual information.
Approach: They propose a multimodal persona-based architecture to generate dialogues for the next panel in comic strips.
Outcome: The proposed paradigm reduces the perplexity score by 10 points over existing models . the novel dataset, ComSet, contains 54K comic strips .
A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)

Copied to clipboard

Challenge: Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener.
Approach: They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen.
Outcome: The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object.
Dialog Intent Structure: A Hierarchical Schema of Linked Dialog Acts (L18-1)

Copied to clipboard

Challenge: a schema for dialog representation captures the pragmatic intents of the conversation independently from any semantic representation.
Approach: They propose a hierarchical and extensible schema for dialog representation . schema captures pragmatic intents of conversation independently from any semantic representation based on semantic content .
Outcome: The proposed schema captures the pragmatic intents of the conversation independently from any semantic representation.
Dialogue Structure Annotation for Multi-Floor Interaction (L18-1)

Copied to clipboard

Challenge: Existing annotation schemes do not address dialogue structure.
Approach: They propose an annotation scheme for meso-level dialogue structure that clusters utterances from multiple participants and floors into units according to realization of an initiator's intent.
Outcome: The proposed annotation scheme is used to annotate a corpus of human-robot interaction dialogues.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations