Action Verb Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of 390 simple actions is based on multimodal data of 12 humans . the dataset is annotated with orthographic transcriptions of utterances and part-of-speech tags .
Approach: They present a multimodal corpus of 12 humans performing 390 simple actions . they also propose an algorithm for segmenting words into utterances and aligning visual information and speech .
Outcome: The presented dataset includes 390 simple actions performed by 12 humans . it includes transcriptions of utterances, part-of-speech tags, lemmata, and hand touches .

Similar Papers

A Corpus of Natural Multimodal Spatial Scene Descriptions (L18-1)

Copied to clipboard

Challenge: Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions.
Approach: They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks.
Outcome: The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces.
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community .
Approach: They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations.
Outcome: The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community .
Linguistic, Kinematic and Gaze Information in Task Descriptions: The LKG-Corpus (2020.lrec-1)

Copied to clipboard

Challenge: linguistic structure of utterances referring to concrete actions may reflect the structure of the sensorimotor processing underlying the same action.
Approach: They present a dataset that integrates linguistic, kinematic and gaze data with an explicit focus on relations between action and language.
Outcome: The proposed dataset integrates linguistic, kinematic and gaze data with an explicit focus on relations between action and language.
Dialogue Act Annotation in a Multimodal Corpus of First Encounter Dialogues (2020.lrec-1)

Copied to clipboard

Challenge: a method used to annotate dialogue acts in a multimodal corpus is described . the annotations allow for analysis of how multimodal signals contribute to the structure and content of the dialogues.
Approach: They propose to annotate dialogue acts in a multimodal corpus of first encounter dialogues . they focus on which dialogue acts often follow each other across speakers and which overlap gestural behaviour .
Outcome: The method used to annotate dialogue acts in a multimodal corpus is described.
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: The goal of the project Multilingual Image Corpus is to provide a large image dataset with annotated objects and object descriptions in 24 languages.
Approach: They propose to provide a large image dataset with annotated objects and object descriptions in 24 languages.
Outcome: The project provides a large image dataset with annotated objects and object descriptions in 24 languages.
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.
A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)

Copied to clipboard

Challenge: Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener.
Approach: They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen.
Outcome: The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object.
A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks (L18-1)

Copied to clipboard

Challenge: Phonetic segmentation is the process of splitting speech into distinct phonetic units . methods for automatic segmentation are not always accurate enough .
Approach: They propose to model phonetic segmentation as close as possible to manual segmentation by recording experts performing a segmentation task.
Outcome: This corpus captures human segmentation behavior by recording experts performing a segmentation task.
In Search of the Lost Arch in Dialogue: A Dependency Dialogue Acts Corpus for Multi-Party Dialogues (2025.findings-acl)

Copied to clipboard

Challenge: Understanding speaker intentions remains a challenge in NLP . a number of corpora annotated using theoretical frameworks of dialogue focus on utterance-level labeling of speaker intent, missing wider context, or the rhetorical structure of a dialogue.
Approach: They propose to annotate a corpus of 33 dialogues and over 9,000 utterance units using the Dependency Dialogue Acts framework.
Outcome: The proposed corpus spans four genres of multi-party conversations from different modalities.
The AICO Multimodal Corpus – Data Collection and Preliminary Analyses (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on human multimodal behaviour in interactions with a human or a robot partner are limited.
Approach: They describe the first explorative research on the AICO Multimodal Corpus, which contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions.
Outcome: The AICO Multimodal Corpus contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations