XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)
Copied to clipboard
| Challenge: | Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. |
| Approach: | They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. |
| Outcome: | The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks. |
Similar Papers
AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation (2023.acl-long)
Copied to clipboard
Rongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Linjun Li, Zhenhui Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiang Yin, Zhou Zhao
| Challenge: | Existing models for speech-to-speech translation suffer from distinct degradation in noisy environments and fail to translate visual speech. |
| Approach: | They propose a text-based audio-visual speech-to-speech translation model that integrates visual information with audio-only data to improve system robustness. |
| Outcome: | The proposed model outperforms models trained on audio-only corpus in two languages . it also improves with low-resource audio-visual data, compared with baselines . |
Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History Interviews (2020.lrec-1)
Copied to clipboard
| Challenge: | Current automatic speech recognition systems show remarkable performance when adequate data is used for training. |
| Approach: | They propose to perform a robust acoustic model adaption to a target domain in a cross-lingual manner. |
| Outcome: | The proposed approach reduces word error rate by more than 30% on German oral history interviews compared to a model trained from scratch on the target domain and 6-7% on same-language out-of-domain training data. |
Hearing Lips in Noise: Universal Viseme-Phoneme Mapping and Transfer for Robust Audio-Visual Speech Recognition (2023.acl-long)
Copied to clipboard
| Challenge: | Existing efforts to improve robustness of audio-visual speech recognition with visual information focus on audio modality . current approaches introduce noise adaptation techniques to improve reliability of AVSR task . |
| Approach: | They propose a visual-invariant modality to strengthen robustness of audio-visual speech recognition (AVSR) it can adapt to any testing noises without dependence on noisy training data, a.k.a., unsupervised noise adaptation. |
| Outcome: | The proposed method outperforms existing state-of-the-arts on visual speech recognition task under various noisy and clean conditions. |
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models. |
| Approach: | They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages. |
| Outcome: | The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data. |
Learning Robust and Multilingual Speech Representations (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Unsupervised speech representation learning has shown success at finding representations that correlate with phonetic structures and improve downstream speech recognition performance. |
| Approach: | They evaluate unsupervised speech representation learning representations by looking at their robustness to domain shifts and their ability to improve recognition performance in many languages. |
| Outcome: | The proposed representations improve the recognition performance in 25 phonetically diverse languages and are robust to domain shifts. |
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)
Copied to clipboard
| Challenge: | Using acoustic data, we develop automatic speech recognition systems for three low resource languages. |
| Approach: | They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework. |
| Outcome: | The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut. |
Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies show that multilingual models outperform monolingual ones. |
| Approach: | They propose a single model that can capture which language is given as input speech . they use a pre-trained model to fine-tune the model so it can recognize the language class as well as the speech with the corresponding language. |
| Outcome: | The proposed model can recognize which language is given as input speech . it can accurately recognize speech in noisy environments, such as crowded restaurants . |
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis. |
| Approach: | They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
| Outcome: | The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. |
Improving Zero-Shot Cross-Lingual Transfer Learning via Robust Training (2021.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained multilingual language encoders do not precisely align words and phrases across languages. |
| Approach: | They propose a learning strategy for training robust models by drawing connections between adversarial examples and failure cases of zero-shot cross-lingual transfer. |
| Outcome: | The proposed model can achieve good performance even if representations of different languages are not aligned well. |
AudioVSR: Enhancing Video Speech Recognition with Audio Data (2024.emnlp-main)
Copied to clipboard
Xiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu, Minjie Hong, Minghui Fang, Shengpeng Ji, Jialong Zuo, Zhiqing Hong, Zhimeng Zhang, Tao Jin
| Challenge: | Recent work has shown poor performance with non-Indo-European languages . previous work primarily utilizes video information to build VSR models . |
| Approach: | They propose a generative model for data inflation that integrates synthetic data with authentic visual data to enhance the VSR model. |
| Outcome: | The proposed model improves on the audio-visual alignment problem in audio-video tasks. |