Deep learning-based end-to-end spoken language identification system for domain-mismatched scenario (2022.lrec-1)
Copied to clipboard
| Challenge: | Domain mismatch is a critical issue when it comes to spoken language identification. |
| Approach: | They evaluated a set of cross-domain language identification trials using a dataset from the Oriental Language Recognition (OLR) Challenge 2021 . |
| Outcome: | The proposed architectures and deep learning strategies have shown good performance in cross-domain speaker verification tasks. |
Similar Papers
Massively Multilingual Adversarial Speech Recognition (N19-1)
Copied to clipboard
| Challenge: | Prior work in multilingual and cross-lingual speech recognition has been limited to a subset of the world's most-spoken languages. |
| Approach: | They propose to use phonemes and phonemes as pretraining objectives to encourage language-independent representations. |
| Outcome: | The proposed model is able to learn language-independent representations of speech using multilingual training. |
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation. |
| Approach: | This tutorial introduces the techniques used in cutting-edge research on speech translation. |
| Outcome: | The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages. |
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)
Copied to clipboard
| Challenge: | Using acoustic data, we develop automatic speech recognition systems for three low resource languages. |
| Approach: | They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework. |
| Outcome: | The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut. |
Cross-lingual Spoken Language Understanding with Regularized Representation Alignment (2020.emnlp-main)
Copied to clipboard
| Challenge: | despite promising results, current cross-lingual models suffer from imperfect cross-linguistic representation alignments between the source and target languages, which makes the performance sub-optimal. |
| Approach: | They propose a regularization approach to align word-level and sentence-level representations across languages without external resources. |
| Outcome: | The proposed model outperforms state-of-the-art models in few-shot and zero-shot scenarios and achieves comparable performance to supervised training with all training data. |
Contrastive and Consistency Learning for Neural Noisy-Channel Model in Spoken Language Understanding (2024.naacl-long)
Copied to clipboard
| Challenge: | End-to-end learning models require large volume of speech data with intent labels . however, models are sensitive to inconsistencies between training and evaluation conditions . |
| Approach: | They propose a module-based approach to learn intent in a noisy-channel model . they correlate error patterns between clean and noisy ASR transcripts . |
| Outcome: | The proposed method outperforms existing methods and improves in noisy environments. |
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models. |
| Approach: | They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages. |
| Outcome: | The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data. |
Cross-domain and Cross-lingual Abusive Language Detection: A Hybrid Approach with Deep Learning and a Multilingual Lexicon (P19-2)
Copied to clipboard
| Challenge: | Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders. |
| Approach: | They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content. |
| Outcome: | The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical. |
Speech Translation and the End-to-End Promise: Taking Stock of Where We Are (2020.acl-main)
Copied to clipboard
| Challenge: | Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach. |
| Approach: | They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity . |
| Outcome: | This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity. |
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Mainstream of automatic speech recognition (ASR) has shifted from pipeline methods to end-to-end (E2E) methods. |
| Approach: | They propose to integrate a pre-trained speech representation model and a large language model (LLM) for automatic speech recognition in an end-to-end manner. |
| Outcome: | The proposed model achieves comparable performance to modern E2E ASR models by utilizing powerful pre-training models with the proposed integrated approach. |
Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake (2024.findings-naacl)
Copied to clipboard
| Challenge: | a recent study has focused on audio deepfake detection (ADD) due to its ability to impersonate and share false, often malicious information. |
| Approach: | They propose to use multilingual speech Pre-Trained models for Audio deepfake detection (ADD) they propose to combine models with existing models to achieve better ADD detection . |
| Outcome: | The proposed models gain knowledge about diverse pitches, accents, and tones, during theirpre-training phase and are more robust to variations. |