Digital Voicing of Silent Speech (2020.emnlp-main)

Copied to clipboard

Challenge: Using electromyography, we can convert silently mouthed words into audible speech . prior work focused on training speech synthesis models from vocalized data .
Approach: They propose a method of training on silent EMG by transferring audio targets from vocalized to silent signals and propose voicing task using muscle sensor measurements.
Outcome: The proposed method greatly improves intelligibility of audio generated from silent EMG compared to baseline that only trains with vocalized data.

Similar Papers

An Improved Model for Voicing Silent Speech (2021.acl-short)

Copied to clipboard

Challenge: Existing models for voicing silent speech use hand-designed features instead of EMG signals.
Approach: They propose to use facial electromyography signals as input instead of hand-designed features to give the model greater flexibility to learn its own features.
Outcome: The proposed model improves state-of-the-art on an open vocabulary intelligibility evaluation by 25.8%.
Non-invasive electromyographic speech neuroprosthesis: a geometric perspective (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for mapping EMG to time-aligned audio limits applicability to patients who can no longer speak.
Approach: They propose a neuromuscular speech interface that translates silently voiced articulations directly into text.
Outcome: The proposed interface can translate silently voiced articulations directly into text without audio transfer . the proposed system could restore communication for patients with speech loss .
emg2speech: synthesizing speech from electromyography using self-supervised speech models (2026.acl-long)

Copied to clipboard

Challenge: a neuromuscular speech interface translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio.
Approach: They propose a neuromuscular speech interface that translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio.
Outcome: The proposed system synthesizes speech without explicit modeling or vocoder training.
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs (2025.acl-short)

Copied to clipboard

Challenge: Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech.
Approach: They propose an EMG adaptor module that maps EMG features to an LLM's input space and achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task.
Outcome: The proposed module achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task.
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)

Copied to clipboard

Challenge: DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Approach: They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Outcome: The proposed method outperforms baseline methods on quantitative and qualitative evaluations.
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously.
Approach: They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has.
Outcome: The proposed method beats text-based systems in terms of perceived emotion and audio quality.
Direct Speech-to-Speech Translation With Discrete Units (2022.acl-long)

Copied to clipboard

Challenge: Existing direct speech-to-speech translation models rely on text generation as an intermediate step.
Approach: They propose a direct speech-to-speech translation model that translates speech from one language to another without relying on intermediate text generation.
Outcome: The proposed model produces 6.7 BLEUs in the Fisher Spanish-English dataset when trained without any text transcripts and with text supervision.
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach (2024.emnlp-main)

Copied to clipboard

Challenge: Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations.
Approach: They show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.
Outcome: Recent advances in speech representation modeling have shown that learning language directly from speech is feasible.
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on speech-to-singing voice conversion (STS) are limited by the scarcity of paired speech-song data and the suboptimal quality of outputs.
Approach: They propose a self-supervised singing voice pre-training model that transforms a speech-to-singing voice into a paired singing voice.
Outcome: The proposed model improves both STS and singing voice synthesis tasks by combining spoken language and a self-supervised singing voice pre-training model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations