Papers by Joakim Gustafson
Evaluating Sampling-based Filler Insertion with Spontaneous TTS (2022.lrec-1)
Copied to clipboard
| Challenge: | Injecting fillers into spoken dialogue systems has a rich history of study . ambiguity of filler occurrence and inter-speaker difference make modeling and evaluation difficult. |
| Approach: | They propose an objective score for filler insertion using sampling-based sampling . they build three models trained on two single-speaker spontaneous corpora and evaluate them with FPP and perceptual tests. |
| Outcome: | The proposed model is useful in analysis but does not correlate well with perceptual MOS. |
Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis (2020.lrec-1)
Copied to clipboard
| Challenge: | Spontaneous speech is unscripted and created on the fly by the speaker, whereas read speech is pre-planned. |
| Approach: | They propose a tool that allows developers to select a varied, representative set of utterances from a spoken genre to be used for evaluation of TTS for a given domain. |
| Outcome: | The proposed tool can be used to evaluate TTS for a given domain using visualisation and tree-based algorithm. |
Revisiting Three Text-to-Speech Synthesis Experiments with a Web-Based Audience Response System (2024.lrec-main)
Copied to clipboard
| Challenge: | Audience Response System (ARS) evaluations are not well understood for text-to-speech synthesis (TTS) evaluation is a key weakness in the field and needs to adapt to be better-suited for this new generation of voices. |
| Approach: | They revisit three published TTS studies and perform an ARS-based evaluation on the stimuli used in each study. |
| Outcome: | The results show that Audience Response System (ARS) is highly useful for evaluating long and continuous stimuli. |
The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTS (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in spontaneous text-to-speech (TTS) have enabled the realistic generation of creaky voice, a voice quality known for its diverse pragmatic and paralinguistic functions. |
| Approach: | They used a creaky voice detection tool and a neural TTS engine to control creaky phonation in a spontaneous speech corpus to investigate the effect of creaky voices on perceived certainty, valence, sarcasm, and turn finality. |
| Outcome: | The proposed model enables the realistic synthesis of creaky voice in perceptual tests without formal training. |
A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)
Copied to clipboard
Dimosthenis Kontogiorgos, Vanya Avramova, Simon Alexanderson, Patrik Jonell, Catharine Oertel, Jonas Beskow, Gabriel Skantze, Joakim Gustafson
| Challenge: | Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener. |
| Approach: | They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen. |
| Outcome: | The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object. |
Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding (2020.lrec-1)
Copied to clipboard
| Challenge: | In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. |
| Approach: | They propose a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture. |
| Outcome: | The proposed method avoids implicit experimenter biases by allowing subjects to instruct each other on the nature of the task: the process of the furniture assembly. |
Crowdsourced Multimodal Corpora Collection Tool (L18-1)
Copied to clipboard
| Challenge: | a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab. |
| Approach: | They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion. |
| Outcome: | The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus. |