Challenge: Conversations and professional interactions are associated with increased risk of SARS-CoV-2 exposure . however, it is unclear to what extent speech properties influence droplets emission .
Approach: They propose to measure velocity and direction of airflow, the number and size of droplets spread during conversation in french.
Outcome: The results will allow future simulation studies to predict the transport, dispersion and evaporation of droplets emitted under different speech conditions.

Similar Papers

Speech Aerodynamics Database, Tools and Visualisation (2022.lrec-1)

Copied to clipboard

Challenge: Aerodynamic processes underlie the characteristics of the acoustic signal of speech sounds.
Approach: a database of aerodynamic processes underlies the characteristics of the acoustic signal of speech sounds . a project was undertaken to obtain data with simultaneous recording of speech acustic signals .
Outcome: the database was designed during an ARC project . it contains recordings of 2 English, 1 Amharic, and 7 French speakers .
Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)

Copied to clipboard

Challenge: Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation .
Approach: They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space.
Outcome: The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis.
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in generative language modeling applied to discrete speech tokens presented a new avenue for text-to-speech (TTS) synthesis.
Approach: They propose to use generative language modeling to generate text-to-speech (TTS) outputs by a discrete token-based model.
Outcome: The proposed model is rated higher in naturalness and context appropriateness in listening tests compared to a conventional TTS.
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models (2025.acl-industry)

Copied to clipboard

Challenge: Text-to-Speech (TTS) training requires extensive and diverse text and speech data.
Approach: They propose a synthetic speech data generation pipeline that generates multilingual, domain-specific datasets for TTS training.
Outcome: The proposed pipeline generates data that is 10–48% more diverse than baseline across various linguistic and phonetic metrics, along with speaker-standardized speech audio while generating approximately 97% correctly normalized text.
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora.
Approach: They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch .
Outcome: The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model.
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
MAD Speech: Measures of Acoustic Diversity of Speech (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in generative spoken language modeling have produced models that produce speech in a wide range of voices, prosody and recording conditions.
Approach: They propose acoustic diversity metrics that measure voice, gender, emotion, accent, background noise and a priori known diversity preferences for each facet.
Outcome: The proposed metrics show that they achieve stronger agreement with diversity than baselines.
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks.
Approach: They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications .
Outcome: The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research.
InteractSpeech: A Speech Dialogue Interaction Corpus for Spoken Dialogue Model (2025.findings-emnlp)

Copied to clipboard

Challenge: Spoken Dialogue models face challenges in handling nuanced interactional phenomena, such as interruptions and backchannels.
Approach: They propose to use a 150-hour English speech interaction dialogue dataset to empower spoken dialogue models with nuanced real-time interaction capabilities.
Outcome: The proposed dataset trains and evaluates a speech understanding model that classifies key interactional events directly from audio.
Design and Development of Speech Corpora for Air Traffic Control Training (L18-1)

Copied to clipboard

Challenge: The current state-of-the-art training procedures involve retired pilots that train as virtual plane pilots and process the spoken prompts to form that can be entered into software that simulates the plane movement on the radar screen.
Approach: They describe the process of creating domain-specific speech corpora containing air traffic control (ATC) communication prompts.
Outcome: The proposed system could be used for training air traffic controllers in the Czech Republic.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations