Challenge: Language model adaptation (LMA) is a promising solution for conversational speech recognition systems.
Approach: They propose to use language model adaptation techniques to adapt language models to conversational speech recognition.
Outcome: The proposed toolkit compares state-of-the-art language model adaptation techniques in conversational speech recognition tasks.

Similar Papers

Spoken Conversational Agents with Large Language Models (2025.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial focuses on the evolution of voice-native LLMs . it reviews the adaptation of text LLM to audio, cross-modal alignment, and joint speech–text training .
Approach: This tutorial examines the evolution of voice-native LLMs in conversational agents . it compares cascaded and voice-based LLM systems to end-to-end retrieval-and vision-grounded systems .
Outcome: This tutorial examines the evolution of voice-native LLMs . it compares the performance of voice assistants to current open-domain agents .
A Survey of Multilingual Models for Automatic Speech Recognition (2022.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, but the majority of the world’s languages do not have usable systems due to the lack of large speech datasets to train these models.
Approach: They propose to use unlabeled speech data to build multilingual ASR models that can be used for improved performance on low-resource languages.
Outcome: The proposed models can be used to improve performance on low-resource languages by using unlabeled speech data.
Session-level Language Modeling for Conversational Speech (D18-1)

Copied to clipboard

Challenge: Xiong et al., 2017) generalizes language models for conversational speech recognition . recurrent neural networks (RNNs) read a list of words sequentially and predict the next word at each position.
Approach: They propose to generalize language models for conversational speech recognition to capture conversation-level phenomena such as adjacency pairs, lexical entrainment, and topical coherence.
Outcome: The proposed model reduces perplexity and improves word error rate over standard models in the conversational telephone speech domain.
Error-preserving Automatic Speech Recognition of Young English Learners’ Language (2024.acl-long)

Copied to clipboard

Challenge: State-of-the-art speech recognition models are often trained on adult read-aloud data by native speakers and do not transfer well to young language learners’ speech.
Approach: They propose to use an automated speech recognition module to train language learners' speaking skills on spontaneous speech by young language learners.
Outcome: The proposed model improves on 85 hours of English audio spoken by Swiss learners and preserves their mistakes.
Discourse on ASR Measurement: Introducing the ARPOCA Assessment Tool (2022.acl-srw)

Copied to clipboard

Challenge: Automated speech recognition (ASR) models are based on a corpus of audio recordings, but are often small or nonexistent for less common languages and dialects.
Approach: This research proposal will develop a semi-automatic acoustic features extraction system that integrates phonetic transcripts with pronunciation dictionaries.
Outcome: The proposed system will be used to improve language recognition and model feedback in less common languages and dialects.
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition (2024.findings-acl)

Copied to clipboard

Challenge: Mainstream of automatic speech recognition (ASR) has shifted from pipeline methods to end-to-end (E2E) methods.
Approach: They propose to integrate a pre-trained speech representation model and a large language model (LLM) for automatic speech recognition in an end-to-end manner.
Outcome: The proposed model achieves comparable performance to modern E2E ASR models by utilizing powerful pre-training models with the proposed integrated approach.
Attention-based Contextual Language Model Adaptation for Speech Recognition (2021.findings-acl)

Copied to clipboard

Challenge: Existing language models do not incorporate utterance level contextual information . however, for some domains like voice assistants, additional context provides a rich input signal .
Approach: They propose a method for training neural speech recognition models on text and contextual data.
Outcome: The proposed model reduces perplexity by 7.0% relative over a standard LM . it also improves perxicity by 2.8% relative to a state-of-the-art model for contextual LM.
A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI (2020.coling-main)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly powerful and more numerous with several options existing as a service.
Approach: They evaluate the most popular automatic speech recognition systems with metrics and experiments designed with these standards in mind.
Outcome: The most popular ASR systems are Microsoft and IBM, and none are suitable for natural spontaneous conversations in real-time.
Personalize Your LLM: Fake it then Align it (2025.findings-naacl)

Copied to clipboard

Challenge: Existing personalization methods require fine-tuning of large language models for each user, rendering them prohibitively expensive for widespread adoption.
Approach: They propose a retrieval-based personalization approach that uses self-generated personal preference data and representation editing to enable quick and cost-effective personalization.
Outcome: The proposed approach outperforms two personalization baselines by 40% on various tasks.
WESR: A Benchmark and Strong Baseline for Word-level Event-Speech Recognition (2026.findings-acl)

Copied to clipboard

Challenge: aaron carroll: the precise localization of non-verbal vocal events remains a critical yet under-explored challenge. carroll says current methods suffer from insufficient task definitions with limited category coverage. carrol: knowing exactly where an event occurred is not enough; knowing exactly what it happened is.
Approach: They propose a taxonomy of 21 vocal events with a new categorization into discrete versus continuous types.
Outcome: The proposed model disentangles ASR errors from event detection while maintaining ASR quality.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations