Challenge: Existing methods for speaker and noise-invariant speech representations use unlabeled audio data to pretrain encoders, generating good representations for downstream tasks like automatic speech recognition (ASR) and speaker identification.
Approach: They propose a domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-in-variant clustering.
Outcome: The proposed method reduces computational resources by 12X compared to state-of-the-art methods while outperforming them in severely distorted speech scenarios.

Similar Papers

XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)

Copied to clipboard

Challenge: Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments.
Approach: They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages.
Outcome: The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks.
Learning Robust and Multilingual Speech Representations (2020.findings-emnlp)

Copied to clipboard

Challenge: Unsupervised speech representation learning has shown success at finding representations that correlate with phonetic structures and improve downstream speech recognition performance.
Approach: They evaluate unsupervised speech representation learning representations by looking at their robustness to domain shifts and their ability to improve recognition performance in many languages.
Outcome: The proposed representations improve the recognition performance in 25 phonetically diverse languages and are robust to domain shifts.
Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History Interviews (2020.lrec-1)

Copied to clipboard

Challenge: Current automatic speech recognition systems show remarkable performance when adequate data is used for training.
Approach: They propose to perform a robust acoustic model adaption to a target domain in a cross-lingual manner.
Outcome: The proposed approach reduces word error rate by more than 30% on German oral history interviews compared to a model trained from scratch on the target domain and 6-7% on same-language out-of-domain training data.
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to learn the transfer from speech to text are unexplored . how to solve the representation discrepancy of speech and text is unexplorable .
Approach: They propose a cooperative acoustic and linguistic representation learning method to fuse and utilize contextual information of speech and text.
Outcome: The proposed method outperforms existing methods on low-resource speech recognition.
MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts (2024.findings-acl)

Copied to clipboard

Challenge: Spoken language understanding (SLU) is a crucial task in task-oriented dialogue systems.
Approach: They propose an ASR-Robust SLU framework based on the mixture-of-experts technique to generate additional transcripts from clean transcripts and use it to weigh the representations of the generated transcripts, ASR transcripts .
Outcome: The proposed framework achieves state-of-the-art on three benchmark SLU datasets.
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis.
Approach: They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Outcome: The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.
Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features (2022.acl-long)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems allow for speech synthesis with unprecedented quality and controllability.
Approach: They use embeddings derived from articulatory vectors rather than phoneme identities to learn phoneme representations that hold across languages.
Outcome: The proposed models fine-tuned on 30 minutes of data in a previously unseen language with language agnostic meta learning.
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication (2024.eacl-short)

Copied to clipboard

Challenge: Recent advances in speech synthesis research have enabled the generation of natural-sounding speech, which has prompted a notable shift in TTS research towards the synthesis of speech in the voices of both seen and unseen speakers.
Approach: They propose a multi-level attention aggregation approach that probes and amplifies various speaker-specific attributes in a hierarchical manner.
Outcome: The proposed model achieves substantial speaker similarity and generalizes to out-of-domain (OOD) cases.
PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric Decoupling (2024.acl-long)

Copied to clipboard

Challenge: Spoken language understanding (SLU) suffers from error propagation from automatic speech recognition (ASR) in actual scenarios.
Approach: They propose a framework which calibrates bias and errors and achieves adaptive-balanced decoupling training by a prototype-based loss model.
Outcome: The proposed framework outperforms existing approaches and achieves state-of-the-art performance on three datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations