Challenge: The production of speech corpora typically involves manual labor to verify and correct the output of automatic transcription/segmentation processes.
Approach: They propose to use Support Vector Machine/Support Vector Regression and Random Forest to predict transcription errors in an annotated speech corpus.
Outcome: The proposed methods can be implemented as free-to-use common language and resources and technology infrastucture web services.

Similar Papers

That doesn’t sound right: Evaluating speech transcription quality in field linguistics corpora (2025.acl-short)

Copied to clipboard

Challenge: Automated speech recognition (ASR) is a popular tool for documenting languages, but field linguists do not have the data to train robust models.
Approach: They propose to use fieldwork data to identify speech transcriptions that may be unsuitable for training ASR models.
Outcome: The proposed measures can be used to identify transcriptions with characteristics common in field data but could be detrimental to ASR training.
Evaluating Workflows for Creating Orthographic Transcripts for Oral Corpora by Transcribing from Scratch or Correcting ASR-Output (2024.lrec-main)

Copied to clipboard

Challenge: Automated speech recognition systems can reduce transcription effort, but few studies have evaluated this potential.
Approach: They compare efforts for manual transcription vs. automatic correction of ASR-output . they use audio recordings from varying settings to create orthographic transcripts .
Outcome: The proposed methods reduce transcription time by 7 times on average for selected data and transcription conventions compared with corrected transcripts . the more complex the primary data, the more time has to be spent on corrections - the paper concludes a similar study could be conducted in 2022 .
RED-ACE: Robust Error Detection for ASR using Confidence Embeddings (2022.emnlp-main)

Copied to clipboard

Challenge: ASR Error Detection (AED) models post-process the output of Automatic Speech Recognition systems, in order to detect transcription errors.
Approach: They propose to use ASR model's word-level confidence scores to combine ASR models with transcribed text to improve AED performance.
Outcome: The proposed models combine the confidence scores and transcribed text into a contextualized representation.
How to Compare Automatically Two Phonological Strings: Application to Intelligibility Measurement in the Case of Atypical Speech (2020.lrec-1)

Copied to clipboard

Challenge: Atypical speech productions must be evaluated with regard to "typical" or "expected" productions . a first test of this method among healthy speakers and patients treated for cancer has proved its validity .
Approach: They propose a method to evaluate "atypical" speech productions based on phonological transcriptions . authors propose to use phonology to compute distances between phonologic forms produced and expected .
Outcome: The proposed method has been validated in a large population of healthy speakers and patients with cancer . it computes distances between phonological forms produced and expected from cost matrices based on features of phonemes .
CEASR: A Corpus for Evaluating Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) systems are increasingly needed for research and practical applications.
Approach: They propose to use public speech corpora to evaluate the quality of automatic speech recognition (ASR) they calculate an average Word Error Rate (WER) per corpus, per system and per corpor-system pair .
Outcome: The proposed corpus evaluates the quality of automatic speech recognition systems using public speech corpora and transcripts generated by state-of-the-art systems.
Consistent Transcription and Translation of Speech (2020.tacl-1)

Copied to clipboard

Challenge: Existing models that translate without transcribing focus on translation quality, while transcription receives less emphasis.
Approach: They propose a method to evaluate consistency and compare different approaches . they propose 'coupled inference' models that feature a coupled inference procedure can achieve strong consistency.
Outcome: The proposed model is poorly suited to the joint transcription/translation task, but is strong enough to train for consistency.
Does Joint Training Really Help Cascaded Speech Translation? (2022.emnlp-main)

Copied to clipboard

Challenge: Currently, in speech translation, the straightforward approach delivers state-of-the-art results, but fundamental challenges such as error propagation remain.
Approach: They propose to combine a cascaded recognition system with a machine translation system to improve cascade speech translation.
Outcome: The proposed methods can improve cascaded speech translation and suggest alternative training methods.
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible.
Approach: They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries.
Outcome: The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country.
A Survey of Confidence Estimation and Calibration in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks in various domains, but they can be unreliable due to factual errors in their generations.
Approach: They summarize recent advances in LLM confidence estimation and calibration and outline their main lessons learned.
Outcome: The proposed methods can be used to assess the reliability of models and to calibrate them across tasks.
Phoneme transcription of endangered languages: an evaluation of recent ASR architectures in the single speaker scenario (2022.findings-acl)

Copied to clipboard

Challenge: Recent work on phonetic transcription is reported to be the bottleneck in endangered languages . however, when a single speaker is involved, small amounts of training are needed .
Approach: They compare automatic speech recognition (ASR) approaches to speaker-dependent phonetic transcription using a common dataset of 11 languages.
Outcome: The proposed system handles morphologically complex languages and writing systems for which no pronunciation dictionary exists.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations