Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History Interviews (2020.lrec-1)
Copied to clipboard
| Challenge: | Current automatic speech recognition systems show remarkable performance when adequate data is used for training. |
| Approach: | They propose to perform a robust acoustic model adaption to a target domain in a cross-lingual manner. |
| Outcome: | The proposed approach reduces word error rate by more than 30% on German oral history interviews compared to a model trained from scratch on the target domain and 6-7% on same-language out-of-domain training data. |
Similar Papers
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models. |
| Approach: | They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages. |
| Outcome: | The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data. |
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)
Copied to clipboard
| Challenge: | Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. |
| Approach: | They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. |
| Outcome: | The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks. |
Massively Multilingual Adversarial Speech Recognition (N19-1)
Copied to clipboard
| Challenge: | Prior work in multilingual and cross-lingual speech recognition has been limited to a subset of the world's most-spoken languages. |
| Approach: | They propose to use phonemes and phonemes as pretraining objectives to encourage language-independent representations. |
| Outcome: | The proposed model is able to learn language-independent representations of speech using multilingual training. |
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling (2022.lrec-1)
Copied to clipboard
| Challenge: | Current state-of-the-art acoustic models can easily comprise more than 100 million parameters. |
| Approach: | They propose to train a gender-balanced, multilingual corpus from 76.000 contributors via Mozilla’s Common Voice project to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation. |
| Outcome: | The proposed model can perform phonetic symbol recognition and validate the quality of the generated phonetic annotation. |
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis. |
| Approach: | They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
| Outcome: | The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
Learning Robust and Multilingual Speech Representations (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Unsupervised speech representation learning has shown success at finding representations that correlate with phonetic structures and improve downstream speech recognition performance. |
| Approach: | They evaluate unsupervised speech representation learning representations by looking at their robustness to domain shifts and their ability to improve recognition performance in many languages. |
| Outcome: | The proposed representations improve the recognition performance in 25 phonetically diverse languages and are robust to domain shifts. |
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition (2024.lrec-main)
Copied to clipboard
Yash Jain, David M. Chan, Pranav Dheram, Aparna Khare, Olabanji Shonibare, Venkatesh Ravichandran, Shalini Ghosh
| Challenge: | Existing methods for pre-training for automatic speech recognition (ASR) focus on single-stage pre-train followed by fine-tuning on downstream task. |
| Approach: | They propose a multi-modal pre-training method that combines unsupervised pre-training with translation-based supervised mid-training. |
| Outcome: | The proposed method improves WERs by 38.45% over baselines on both Librispeech and SUPERB. |
Cross-lingual Transfer Learning with Data Selection for Large-Scale Spoken Language Understanding (D19-1)
Copied to clipboard
| Challenge: | Existing approaches to improve cross-lingual transfer learning on spoken language are pre-train on all available supervised data from another language. |
| Approach: | They propose a language model based source-language data selection method for cross-lingual transfer learning in spoken language understanding. |
| Outcome: | The proposed method reduces training time and improves model performance on spoken language understanding. |
MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery (2026.acl-long)
Copied to clipboard
| Challenge: | MauBERT models learn from multilingual data to predict articulatory features or phones, resulting in language-independent phonetic representations. |
| Approach: | They introduce a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. |
| Outcome: | The proposed model can predict phonetic features in 55 languages with minimal fine-tuning (10 hours of speech) it is more context-invariant than state-of-the-art models and adapts to unseen languages and casual speech with minimal self-supervised fine- tuning (10 hours) |
Cross-lingual Spoken Language Understanding with Regularized Representation Alignment (2020.emnlp-main)
Copied to clipboard
| Challenge: | despite promising results, current cross-lingual models suffer from imperfect cross-linguistic representation alignments between the source and target languages, which makes the performance sub-optimal. |
| Approach: | They propose a regularization approach to align word-level and sentence-level representations across languages without external resources. |
| Outcome: | The proposed model outperforms state-of-the-art models in few-shot and zero-shot scenarios and achieves comparable performance to supervised training with all training data. |