An Improved Model for Voicing Silent Speech (2021.acl-short)

Copied to clipboard

Challenge: Existing models for voicing silent speech use hand-designed features instead of EMG signals.
Approach: They propose to use facial electromyography signals as input instead of hand-designed features to give the model greater flexibility to learn its own features.
Outcome: The proposed model improves state-of-the-art on an open vocabulary intelligibility evaluation by 25.8%.

Similar Papers

Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach (2024.emnlp-main)

Copied to clipboard

Challenge: Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations.
Approach: They show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.
Outcome: Recent advances in speech representation modeling have shown that learning language directly from speech is feasible.
emg2speech: synthesizing speech from electromyography using self-supervised speech models (2026.acl-long)

Copied to clipboard

Challenge: a neuromuscular speech interface translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio.
Approach: They propose a neuromuscular speech interface that translates electromyographic (EMG) signals recorded from orofacial muscles during speech articulation directly into audio.
Outcome: The proposed system synthesizes speech without explicit modeling or vocoder training.
Non-invasive electromyographic speech neuroprosthesis: a geometric perspective (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for mapping EMG to time-aligned audio limits applicability to patients who can no longer speak.
Approach: They propose a neuromuscular speech interface that translates silently voiced articulations directly into text.
Outcome: The proposed interface can translate silently voiced articulations directly into text without audio transfer . the proposed system could restore communication for patients with speech loss .
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs (2025.acl-short)

Copied to clipboard

Challenge: Unvoiced electromyography (EMG) is an effective communication tool for individuals unable to produce vocal speech.
Approach: They propose an EMG adaptor module that maps EMG features to an LLM's input space and achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task.
Outcome: The proposed module achieves an average word error rate of 0.49 on a closed-vocabulary unvoiced EMG-to-text task.
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)

Copied to clipboard

Challenge: DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Approach: They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Outcome: The proposed method outperforms baseline methods on quantitative and qualitative evaluations.
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning (2024.emnlp-main)

Copied to clipboard

Challenge: EH-MAM is a self-supervised learning approach for speech representation learning . prior methods used random masking schemes to learn speech representations .
Approach: They propose a self-supervised approach that automatically selects hard regions during SSL training and introduces them to the model for reconstruction.
Outcome: The proposed approach outperforms state-of-the-art models across low-resource speech recognition and SUPERB benchmarks by 5%-10%.
Simple and Effective Unsupervised Speech Translation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to train speech models without labeled data are limited for most languages.
Approach: They propose a pipeline approach to build speech translation systems without labeled data by leveraging recent advances in unsupervised speech recognition, machine translation and speech synthesis.
Outcome: The proposed approach outperforms the state-of-the-art in unsupervised speech recognition by 3.2 BLEU on the Libri-Trans benchmark and the best supervised end-to-end models from only two years ago by an average of 5.0 BLUE over five X-En directions.
From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation model (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models that map variable acoustic inputs into appropriate articulatory movements without explicit instruction are inadequate for infants.
Approach: They propose a model that maps acoustic inputs into articulatory movements without explicit instruction for infants.
Outcome: The proposed model outperforms MFCC features in both single- and multi-speaker settings and provides optimal representations for articulatory learning.
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis.
Approach: They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Outcome: The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations