| Challenge: | Using electromyography, we can convert silently mouthed words into audible speech . prior work focused on training speech synthesis models from vocalized data . |
| Approach: | They propose a method of training on silent EMG by transferring audio targets from vocalized to silent signals and propose voicing task using muscle sensor measurements. |
| Outcome: | The proposed method greatly improves intelligibility of audio generated from silent EMG compared to baseline that only trains with vocalized data. |
Similar Papers
An Improved Model for Voicing Silent Speech (2021.acl-short)
Copied to clipboard
| Challenge: | Existing models for voicing silent speech use hand-designed features instead of EMG signals. |
| Approach: | They propose to use facial electromyography signals as input instead of hand-designed features to give the model greater flexibility to learn its own features. |
| Outcome: | The proposed model improves state-of-the-art on an open vocabulary intelligibility evaluation by 25.8%. |
Non-invasive electromyographic speech neuroprosthesis: a geometric perspective (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for mapping EMG to time-aligned audio limits applicability to patients who can no longer speak. |
| Approach: | They propose a neuromuscular speech interface that translates silently voiced articulations directly into text. |
| Outcome: | The proposed interface can translate silently voiced articulations directly into text without audio transfer . the proposed system could restore communication for patients with speech loss . |
emg2speech: synthesizing speech from electromyography using self-supervised speech models (2026.acl-long)
Copied to clipboard
| Challenge: | a neuromuscular speech interface translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio. |
| Approach: | They propose a neuromuscular speech interface that translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio. |
| Outcome: | The proposed system synthesizes speech without explicit modeling or vocoder training. |
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs (2025.acl-short)
Copied to clipboard
| Challenge: | Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech. |
| Approach: | They propose an EMG adaptor module that maps EMG features to an LLM's input space and achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task. |
| Outcome: | The proposed module achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task. |
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)
Copied to clipboard
| Challenge: | DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. |
| Approach: | They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. |
| Outcome: | The proposed method outperforms baseline methods on quantitative and qualitative evaluations. |
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)
Copied to clipboard
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, Emmanuel Dupoux
| Challenge: | Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language. |
| Approach: | They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels. |
| Outcome: | The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation. |
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)
Copied to clipboard
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu Anh Nguyen, Morgan Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi
| Challenge: | Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously. |
| Approach: | They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has. |
| Outcome: | The proposed method beats text-based systems in terms of perceived emotion and audio quality. |
Direct Speech-to-Speech Translation With Discrete Units (2022.acl-long)
Copied to clipboard
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, Wei-Ning Hsu
| Challenge: | Existing direct speech-to-speech translation models rely on text generation as an intermediate step. |
| Approach: | They propose a direct speech-to-speech translation model that translates speech from one language to another without relying on intermediate text generation. |
| Outcome: | The proposed model produces 6.7 BLEUs in the Fisher Spanish-English dataset when trained without any text transcripts and with text supervision. |
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach (2024.emnlp-main)
Copied to clipboard
| Challenge: | Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. |
| Approach: | They show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data. |
| Outcome: | Recent advances in speech representation modeling have shown that learning language directly from speech is feasible. |
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on speech-to-singing voice conversion (STS) are limited by the scarcity of paired speech-song data and the suboptimal quality of outputs. |
| Approach: | They propose a self-supervised singing voice pre-training model that transforms a speech-to-singing voice into a paired singing voice. |
| Outcome: | The proposed model improves both STS and singing voice synthesis tasks by combining spoken language and a self-supervised singing voice pre-training model. |