Challenge: Existing studies on hand gestures from video-recorded speeches have not identified them.
Approach: They annotated and analysed hand gestures produced by Barack Obama . they trained machine learning algorithms to classify the semiotic type of hand gesture .
Outcome: The proposed method can be used to classify hand gestures on video-recorded speeches and in advanced multimodal interactive systems.

Similar Papers

Adding Gesture, Posture and Facial Displays to the PoliModal Corpus of Political Interviews (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of face-to-face interviews is annotated with facial displays, hand gestures and body posture.
Approach: They introduce a multimodal corpus on top of transcribed face-to-face interviews that presents the annotation of facial displays, hand gestures and body posture.
Outcome: The proposed corpus is extracted from a larger corpus of 56 face-to-face interviews (14 hours) the annotations include facial displays, hand gestures and body posture.
Creating a Corpus of Gestures and Predicting the Audience Response based on Gestures in Speeches of Donald Trump (2020.lrec-1)

Copied to clipboard

Challenge: a study aims to explore the role of speech pauses and gestures alone as predictors of audience reaction without other types of speech information.
Approach: They analyze two speeches by Barack Obama and use them to predict audience reaction . they find that long pauses and co-speech gestures alone predict audience response .
Outcome: The proposed models can predict audience reaction without other types of speech information.
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
The POTUS Corpus, a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents (2020.lrec-1)

Copied to clipboard

Challenge: Embodied Conversational Agents (ECAs) are used to generate socially believable agents.
Approach: They propose to use audio-video files of political addresses to generate a corpus of socially believable agents which can be annotated by external observers.
Outcome: The proposed corpus analyzes audio-video files of political addresses to the american people and provides the same speeches given by a virtual agent named Rodrigue.
LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures (2024.acl-long)

Copied to clipboard

Challenge: despite advances in the generation of realistic human gestures, the process often includes unintended, meaningless, or non-realistic gestures.
Approach: They propose a framework that leverages large language models to generate human gestures . the primary stage employs a transformer-based auto-encoder network to encode human gesture into discrete symbols .
Outcome: The proposed framework has demonstrated state-of-the-art performance on public TED and TED-Expressive datasets.
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies (2024.lrec-main)

Copied to clipboard

Challenge: a small fraction of the languages currently covered by speech technologies are mainly spoken in English.
Approach: They present an annotation toolkit that detects when a person speaks on the scene and the corresponding transcription.
Outcome: The proposed toolkit can speed up the annotation process by up to four times . it can be used in Spanish, and is available on github.
Machine-Aided Annotation for Fine-Grained Proposition Types in Argumentation (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 2016 debates and commentary contains 4,648 argumentative propositions annotated with fine-grained proposition types.
Approach: They propose a machine learning-human workflow for annotating for four complex proposition types . they demonstrate with preliminary analysis of rhetorical strategies and structure in presidential debates .
Outcome: The proposed method can be used by technical researchers seeking more nuanced representations of argument . it can also be used to analyze rhetorical strategies and structure in presidential debates .
Development of an Annotated Multimodal Dataset for the Investigation of Classification and Summarisation of Presentations using High-Level Paralinguistic Features (L18-1)

Copied to clipboard

Challenge: Existing summarisation methods take no account of multimodal high-level paralinguistic features which form part of audio-visual presentations.
Approach: They propose to use audiovisual recordings to extract paralinguistic features from audio recordings . they use manual annotations to help users find relevant material .
Outcome: The proposed method can identify the most important or emphasised material within a presentation.
Towards Understanding the Relation between Gestures and Language (2022.coling-1)

Copied to clipboard

Challenge: a new study explores the relationship between gestures and language . we use contrastive learning to learn gesture embeddings .
Approach: They adapt a semi-supervised multimodal model to learn gesture embeddings using Ted talks . they show gestures are predictive of the native language of the speaker .
Outcome: The proposed model learns gesture embeddings from a multimodal dataset . it shows that gesture embeds are predictive of the native language of the speaker .
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)

Copied to clipboard

Challenge: Using the corpus, we study the characteristics of interpreters' work and train machine translation systems.
Approach: They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work.
Outcome: The proposed corpus can be used for teaching interpreters and to train machine translation systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations