R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for speaker and noise-invariant speech representations use unlabeled audio data to pretrain encoders, generating good representations for downstream tasks like automatic speech recognition (ASR) and speaker identification. |
| Approach: | They propose a domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-in-variant clustering. |
| Outcome: | The proposed method reduces computational resources by 12X compared to state-of-the-art methods while outperforming them in severely distorted speech scenarios. |
Similar Papers
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)
Copied to clipboard
| Challenge: | Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. |
| Approach: | They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. |
| Outcome: | The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks. |
Learning Robust and Multilingual Speech Representations (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Unsupervised speech representation learning has shown success at finding representations that correlate with phonetic structures and improve downstream speech recognition performance. |
| Approach: | They evaluate unsupervised speech representation learning representations by looking at their robustness to domain shifts and their ability to improve recognition performance in many languages. |
| Outcome: | The proposed representations improve the recognition performance in 25 phonetically diverse languages and are robust to domain shifts. |
Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History Interviews (2020.lrec-1)
Copied to clipboard
| Challenge: | Current automatic speech recognition systems show remarkable performance when adequate data is used for training. |
| Approach: | They propose to perform a robust acoustic model adaption to a target domain in a cross-lingual manner. |
| Outcome: | The proposed approach reduces word error rate by more than 30% on German oral history interviews compared to a model trained from scratch on the target domain and 6-7% on same-language out-of-domain training data. |
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to learn the transfer from speech to text are unexplored . how to solve the representation discrepancy of speech and text is unexplorable . |
| Approach: | They propose a cooperative acoustic and linguistic representation learning method to fuse and utilize contextual information of speech and text. |
| Outcome: | The proposed method outperforms existing methods on low-resource speech recognition. |
MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts (2024.findings-acl)
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) is a crucial task in task-oriented dialogue systems. |
| Approach: | They propose an ASR-Robust SLU framework based on the mixture-of-experts technique to generate additional transcripts from clean transcripts and use it to weigh the representations of the generated transcripts, ASR transcripts . |
| Outcome: | The proposed framework achieves state-of-the-art on three benchmark SLU datasets. |
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis. |
| Approach: | They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
| Outcome: | The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)
Copied to clipboard
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, Emmanuel Dupoux
| Challenge: | Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language. |
| Approach: | They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels. |
| Outcome: | The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation. |
Language-Agnostic Meta-Learning for Low-Resource Text-to-Speech with Articulatory Features (2022.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-speech systems allow for speech synthesis with unprecedented quality and controllability. |
| Approach: | They use embeddings derived from articulatory vectors rather than phoneme identities to learn phoneme representations that hold across languages. |
| Outcome: | The proposed models fine-tuned on 30 minutes of data in a previously unseen language with language agnostic meta learning. |
Multi-Level Attention Aggregation for Language-Agnostic Speaker Replication (2024.eacl-short)
Copied to clipboard
| Challenge: | Recent advances in speech synthesis research have enabled the generation of natural-sounding speech, which has prompted a notable shift in TTS research towards the synthesis of speech in the voices of both seen and unseen speakers. |
| Approach: | They propose a multi-level attention aggregation approach that probes and amplifies various speaker-specific attributes in a hierarchical manner. |
| Outcome: | The proposed model achieves substantial speaker similarity and generalizes to out-of-domain (OOD) cases. |
PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric Decoupling (2024.acl-long)
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) suffers from error propagation from automatic speech recognition (ASR) in actual scenarios. |
| Approach: | They propose a framework which calibrates bias and errors and achieves adaptive-balanced decoupling training by a prototype-based loss model. |
| Outcome: | The proposed framework outperforms existing approaches and achieves state-of-the-art performance on three datasets. |