Papers by Joakim Gustafson

7 papers
Evaluating Sampling-based Filler Insertion with Spontaneous TTS (2022.lrec-1)

Copied to clipboard

Challenge: Injecting fillers into spoken dialogue systems has a rich history of study . ambiguity of filler occurrence and inter-speaker difference make modeling and evaluation difficult.
Approach: They propose an objective score for filler insertion using sampling-based sampling . they build three models trained on two single-speaker spontaneous corpora and evaluate them with FPP and perceptual tests.
Outcome: The proposed model is useful in analysis but does not correlate well with perceptual MOS.
Augmented Prompt Selection for Evaluation of Spontaneous Speech Synthesis (2020.lrec-1)

Copied to clipboard

Challenge: Spontaneous speech is unscripted and created on the fly by the speaker, whereas read speech is pre-planned.
Approach: They propose a tool that allows developers to select a varied, representative set of utterances from a spoken genre to be used for evaluation of TTS for a given domain.
Outcome: The proposed tool can be used to evaluate TTS for a given domain using visualisation and tree-based algorithm.
Revisiting Three Text-to-Speech Synthesis Experiments with a Web-Based Audience Response System (2024.lrec-main)

Copied to clipboard

Challenge: Audience Response System (ARS) evaluations are not well understood for text-to-speech synthesis (TTS) evaluation is a key weakness in the field and needs to adapt to be better-suited for this new generation of voices.
Approach: They revisit three published TTS studies and perform an ARS-based evaluation on the stimuli used in each study.
Outcome: The results show that Audience Response System (ARS) is highly useful for evaluating long and continuous stimuli.
The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTS (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in spontaneous text-to-speech (TTS) have enabled the realistic generation of creaky voice, a voice quality known for its diverse pragmatic and paralinguistic functions.
Approach: They used a creaky voice detection tool and a neural TTS engine to control creaky phonation in a spontaneous speech corpus to investigate the effect of creaky voices on perceived certainty, valence, sarcasm, and turn finality.
Outcome: The proposed model enables the realistic synthesis of creaky voice in perceptual tests without formal training.
A Multimodal Corpus for Mutual Gaze and Joint Attention in Multiparty Situated Interaction (L18-1)

Copied to clipboard

Challenge: Using a multisensory setup, we capture speech, eye gaze and gesture data and investigate four different types of social gaze: referential gaze, joint attention, mutual gaze and gaze aversion by both perspectives of a speaker and a listener.
Approach: They present a corpus of situated interaction where participants collaborated on moving virtual objects on a large touch screen.
Outcome: The authors capture speech, eye gaze and gesture data using a multisensory setup and analysed the groups' referential eye-gaze with respect to the referent object.
Chinese Whispers: A Multimodal Dataset for Embodied Language Grounding (2020.lrec-1)

Copied to clipboard

Challenge: In this paper, we introduce a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture.
Approach: They propose a multimodal dataset in which subjects are instructing each other how to assemble IKEA furniture.
Outcome: The proposed method avoids implicit experimenter biases by allowing subjects to instruct each other on the nature of the task: the process of the furniture assembly.
Crowdsourced Multimodal Corpora Collection Tool (L18-1)

Copied to clipboard

Challenge: a crowd-sourced corpora recording method has several disadvantages, including the cost of staff, equipment and time spent recording in-lab.
Approach: They propose to use a crowd-sourced data collection tool to gather controlled multimodal data of people in a rapid and scalable fashion.
Outcome: The proposed tool will allow researchers to quickly gather large amounts of multimodal data spanning a wide demographic range and create their own multimodal corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations