Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.

Similar Papers

A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)

Copied to clipboard

Challenge: Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener.
Approach: They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen.
Outcome: The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object.
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.
MULTICOLLAB: A Multimodal Corpus of Dialogues for Analyzing Collaboration and Frustration in Language (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to study complex emotions when a speaker collaborates with a partner are limited.
Approach: They propose to fuse a multimodal dialogue resource with transcribed speech and eye gaze data to create a highly multimodal corpus.
Outcome: The proposed model improves classification accuracy by 21% over baseline using sensor and speech data in 4.5 seconds.
The AICO Multimodal Corpus – Data Collection and Preliminary Analyses (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on human multimodal behaviour in interactions with a human or a robot partner are limited.
Approach: They describe the first explorative research on the AICO Multimodal Corpus, which contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions.
Outcome: The AICO Multimodal Corpus contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions.
The Niki and Julie Corpus: Collaborative Multimodal Dialogues between Humans, Robots, and Virtual Agents (L18-1)

Copied to clipboard

Challenge: Niki and Julie corpus contains more than 600 dialogues between humans and robots . corpus includes audio and video recordings, results of ranking tasks, questionnaire responses .
Approach: the corpus contains more than 600 dialogues between human participants and a robot . the dialogues are part of a collaborative item-ranking task designed to measure influence .
Outcome: the corpus contains more than 600 dialogues between human participants and a robot or virtual agent . the dialogues contain conversational errors by the robot, which simulates typical of modern automated agents .
Crowdsourced Multimodal Corpora Collection Tool (L18-1)

Copied to clipboard

Challenge: a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab.
Approach: They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion.
Outcome: The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus.
Dialogue Act Annotation in a Multimodal Corpus of First Encounter Dialogues (2020.lrec-1)

Copied to clipboard

Challenge: a method used to annotate dialogue acts in a multimodal corpus is described . the annotations allow for analysis of how multimodal signals contribute to the structure and content of the dialogues.
Approach: They propose to annotate dialogue acts in a multimodal corpus of first encounter dialogues . they focus on which dialogue acts often follow each other across speakers and which overlap gestural behaviour .
Outcome: The method used to annotate dialogue acts in a multimodal corpus is described.
Encoding Gesture in Multimodal Dialogue: Creating a Corpus of Multimodal AMR (2024.lrec-main)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) was designed to represent sentence meaning in English text, but recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture.
Approach: They propose to annotate a multimodal (speech and gesture) AMR corpus in a task-based setting and capture coreference relationships across modalities.
Outcome: The proposed corpus captures coreference relationships across modalities, enabling fine-grained analysis of how gesture and natural language interact.
RoomReader: A Multimodal Corpus of Online Multiparty Conversational Interactions (2022.lrec-1)

Copied to clipboard

Challenge: The corpus of multimodal, multiparty conversational interactions explored in RoomReader can be used to study a wide range of phenomena in online multimodal interaction.
Approach: They propose to use RoomReader to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments.
Outcome: The corpus was developed within the wider RoomReader Project to explore multimodal cues of conversational engagement and behavioural aspects of collaborative interaction in online environments.
Multimodal Large Language Models for Human-AI Interaction: Foundations, Agents, and Inclusive Applications (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Approach: This tutorial presents foundations, agentic capabilities, and inclusive applications of multimodal large language models.
Outcome: This tutorial covers foundations, agentic capabilities, and inclusive applications of multimodal large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations