Challenge: Existing technologies for speech recognition and speech synthesis focus on non-verbal content and paralinguistic information.
Approach: They propose to construct a multimodal affective conversational corpus based on TV dramas . their data contain parallel English-French languages in lexical, acoustic, and facial features .
Outcome: The proposed corpus can be used to assess speech recognition, speech recognition and synthesis, linguistic, and paralinguistic speech-to-speech translation and multimodal dialog systems.

Similar Papers

MUCH: A Multimodal Corpus Construction for Conversational Humor Recognition Based on Chinese Sitcom (2024.lrec-main)

Copied to clipboard

Challenge: Existing multimodal corpora for conversational humor are coarse-grained and insufficient to support the conversational comprehension task.
Approach: They constructed a multimodal humor corpus based on a Chinese sitcom and used both unimodal and multimodal methods to test the corpus.
Outcome: The proposed method outperforms unimodal and multimodal methods in the evaluation of a Chinese sitcom for conversational humor recognition.
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations (P19-1)

Copied to clipboard

Challenge: Emotion recognition in conversations has gained popularity due to its potential applications. Until now, a large multimodal multi-party emotional conversational database containing more than two speakers per dialogue was missing.
Approach: They propose to extend and enhance EmotionLines by combining 13,000 utterances from Friends dialogues with emotion and sentiment labels.
Outcome: The proposed dataset contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends.
Construction and Analysis of a Multimodal Chat-talk Corpus for Dialog Systems Considering Interpersonal Closeness (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal dialog corpus is needed to accelerate research on dialog systems that can handle social signals and verbal information.
Approach: They construct a multimodal dialog corpus focusing on the relationship between speakers and 19 pairs of participants.
Outcome: The proposed system is based on a multimodal dialog corpus of 19,303 utterances (10 hours) from 19 pairs of participants.
EmotionLines: An Emotion Corpus of Multi-Party Conversations (L18-1)

Copied to clipboard

Challenge: Emotion is a critical characteristic to distinguish people from machines.
Approach: They propose a dataset with emotions labeling on all utterances in each dialogue . they use Friends TV scripts and Facebook messenger dialogues to collect the data .
Outcome: The proposed dataset is the first with emotions labeling on all utterances in each dialogue based on their textual content.
DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on multi-party dialogue discourse parsing focus on textual modality and two-party dialog . et al., 2016) focused on text-based discourse parses, ignoring the complexity and richness of multimodal interactions in real-world scenarios.
Approach: They construct the first publicly available English multimodal dataset for multi-party dialogue discourse parsing based on American TV dramas.
Outcome: The proposed dataset contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, covering rich multi-party interaction scenarios.
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups.
Approach: They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks.
Outcome: The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants.
M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue Database (2022.acl-long)

Copied to clipboard

Challenge: Existing data resources to support multimodal affective analysis in dialogues are limited in scale and diversity.
Approach: They propose a multimodal multi-scene multi-label Emotional Dialogue dataset, M3ED, which contains 990 dyadic emotional dialogues from 56 different TV series.
Outcome: The proposed dataset contains 990 dyadic emotional dialogues from 56 different TV series, a total of 9,082 turns and 24,449 utterances.
A Large Scale Speech Sentiment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpus for sentiment analysis uses text inputs, but voice inputs are becoming more important as smart assistants and mobile voice control become more prevalent.
Approach: They propose to extend the Switchboard-1 Telephone Speech Corpus by adding sentiment labels from 3 different human annotators for every transcript segment.
Outcome: The proposed corpus contains 49500 labeled speech segments covering 140 hours of audio.
EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text (L18-1)

Copied to clipboard

Challenge: Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters.
Approach: They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators.
Outcome: The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators.
SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis. (L18-1)

Copied to clipboard

Challenge: a French audiobooks corpus contains 87 hours of good audio quality speech . audiobooks provide mono-genre and multi-speaker speech whereas audiobooks usually provide a few hours of mono- and multispeakers .
Approach: They present an expressive French audiobooks corpus containing eighty seven hours of speech . the corpus is annotated automatically and provides information as phone labels, phone boundaries, syllables, words or morpho-syntactic tagging.
Outcome: The proposed corpus contains 87 hours of speech recorded by a single speaker . the data will allow developing models to better control expressiveness in speech synthesis .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations