| Challenge: | Existing models for voicing silent speech use hand-designed features instead of EMG signals. |
| Approach: | They propose to use facial electromyography signals as input instead of hand-designed features to give the model greater flexibility to learn its own features. |
| Outcome: | The proposed model improves state-of-the-art on an open vocabulary intelligibility evaluation by 25.8%. |
Similar Papers
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach (2024.emnlp-main)
Copied to clipboard
| Challenge: | Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. |
| Approach: | They show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data. |
| Outcome: | Recent advances in speech representation modeling have shown that learning language directly from speech is feasible. |
emg2speech: synthesizing speech from electromyography using self-supervised speech models (2026.acl-long)
Copied to clipboard
| Challenge: | a neuromuscular speech interface translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio. |
| Approach: | They propose a neuromuscular speech interface that translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio. |
| Outcome: | The proposed system synthesizes speech without explicit modeling or vocoder training. |
Non-invasive electromyographic speech neuroprosthesis: a geometric perspective (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for mapping EMG to time-aligned audio limits applicability to patients who can no longer speak. |
| Approach: | They propose a neuromuscular speech interface that translates silently voiced articulations directly into text. |
| Outcome: | The proposed interface can translate silently voiced articulations directly into text without audio transfer . the proposed system could restore communication for patients with speech loss . |
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs (2025.acl-short)
Copied to clipboard
| Challenge: | Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech. |
| Approach: | They propose an EMG adaptor module that maps EMG features to an LLM's input space and achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task. |
| Outcome: | The proposed module achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task. |
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)
Copied to clipboard
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, Emmanuel Dupoux
| Challenge: | Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language. |
| Approach: | They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels. |
| Outcome: | The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation. |
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)
Copied to clipboard
| Challenge: | DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. |
| Approach: | They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. |
| Outcome: | The proposed method outperforms baseline methods on quantitative and qualitative evaluations. |
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning (2024.emnlp-main)
Copied to clipboard
| Challenge: | EH-MAM is a self-supervised learning approach for speech representation learning . prior methods used random masking schemes to learn speech representations . |
| Approach: | They propose a self-supervised approach that automatically selects hard regions during SSL training and introduces them to the model for reconstruction. |
| Outcome: | The proposed approach outperforms state-of-the-art models across low-resource speech recognition and SUPERB benchmarks by 5%-10%. |
Simple and Effective Unsupervised Speech Translation (2023.acl-long)
Copied to clipboard
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino
| Challenge: | Existing methods to train speech models without labeled data are limited for most languages. |
| Approach: | They propose a pipeline approach to build speech translation systems without labeled data by leveraging recent advances in unsupervised speech recognition, machine translation and speech synthesis. |
| Outcome: | The proposed approach outperforms the state-of-the-art in unsupervised speech recognition by 3.2 BLEU on the Libri-Trans benchmark and the best supervised end-to-end models from only two years ago by an average of 5.0 BLUE over five X-En directions. |
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing models that map variable acoustic inputs into appropriate articulatory movements without explicit instruction are inadequate for infants. |
| Approach: | They propose a model that maps acoustic inputs into articulatory movements without explicit instruction for infants. |
| Outcome: | The proposed model outperforms MFCC features in both single- and multi-speaker settings and provides optimal representations for articulatory learning. |
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis. |
| Approach: | They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
| Outcome: | The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |