Challenge: Recent studies show that multilingual models outperform monolingual ones.
Approach: They propose a single model that can capture which language is given as input speech . they use a pre-trained model to fine-tune the model so it can recognize the language class as well as the speech with the corresponding language.
Outcome: The proposed model can recognize which language is given as input speech . it can accurately recognize speech in noisy environments, such as crowded restaurants .

Similar Papers

A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: AVLM integrates full-face visual cues into a pre-trained expressive speech model.
Approach: They propose an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model.
Outcome: The proposed model incorporates full-face visual cues into a pre-trained expressive speech model.
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception (2024.acl-long)

Copied to clipboard

Challenge: Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments.
Approach: They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages.
Outcome: The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks.
Massively Multilingual Adversarial Speech Recognition (N19-1)

Copied to clipboard

Challenge: Prior work in multilingual and cross-lingual speech recognition has been limited to a subset of the world's most-spoken languages.
Approach: They propose to use phonemes and phonemes as pretraining objectives to encourage language-independent representations.
Outcome: The proposed model is able to learn language-independent representations of speech using multilingual training.
Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to synthesize speech for low-resource languages require a substantial amount of source language corpora to generate the linguistic knowledge that can be reused for speech synthesis.
Approach: They propose a method that extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Outcome: The proposed method extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre.
Multi-Staged Cross-Lingual Acoustic Model Adaption for Robust Speech Recognition in Real-World Applications - A Case Study on German Oral History Interviews (2020.lrec-1)

Copied to clipboard

Challenge: Current automatic speech recognition systems show remarkable performance when adequate data is used for training.
Approach: They propose to perform a robust acoustic model adaption to a target domain in a cross-lingual manner.
Outcome: The proposed approach reduces word error rate by more than 30% on German oral history interviews compared to a model trained from scratch on the target domain and 6-7% on same-language out-of-domain training data.
Common Phone: A Multilingual Dataset for Robust Acoustic Modelling (2022.lrec-1)

Copied to clipboard

Challenge: Current state-of-the-art acoustic models can easily comprise more than 100 million parameters.
Approach: They propose to train a gender-balanced, multilingual corpus from 76.000 contributors via Mozilla’s Common Voice project to perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
Outcome: The proposed model can perform phonetic symbol recognition and validate the quality of the generated phonetic annotation.
A Tour of Explicit Multilingual Semantics: Word Sense Disambiguation, Semantic Role Labeling and Semantic Parsing (2022.aacl-tutorials)

Copied to clipboard

Challenge: a recent advent of pretrained language models has sparked a revolution in NLP . but, there are still questions about whether current approaches capture explicit, symbolic meaning . this tutorial will review efforts to tackle three key open problems in lexical and sentence-level semantics .
Approach: This tutorial reviews recent efforts to shed light on meaning in NLP . it will focus on three key open problems in lexical and sentence-level semantics .
Outcome: This tutorial reviews recent efforts to shed light on meaning in NLP . it focuses on three key open problems in lexical and sentence-level semantics .
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)

Copied to clipboard

Challenge: Using acoustic data, we develop automatic speech recognition systems for three low resource languages.
Approach: They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework.
Outcome: The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut.
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives (2024.findings-acl)

Copied to clipboard

Challenge: Existing video-language understanding systems with human-like senses can mimic both our linguistic medium and visual environment with temporal dynamics.
Approach: They propose to develop video-language understanding systems with human-like senses . they summarize their methods and highlight challenges associated with them .
Outcome: The proposed models perform well in a variety of tasks and domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations