Challenge: Pre-trained Large Acoustic Models have been shown to improve performance in spoken languages . however, their potential for novel under-resourced languages is not fully known .
Approach: They propose to use pre-trained Large Acoustic Models to document under-resourced languages . they use scripts from languages that hold a prominent presence in the geographical regions .
Outcome: The proposed model can document under-resourced languages in the electronic domain . the model can be used to document languages with a written script .

Similar Papers

An (unhelpful) guide to selecting the best ASR architecture for your under-resourced language (2023.acl-short)

Copied to clipboard

Challenge: English ASR now has word error rates comparable to that of human transcriptionists, but only for the handful of the world's 7000 languages with abundant training resources.
Approach: They propose to use four of the most popular ASR toolkits to train ASR models for eleven languages with limited ASR training resources: eleven widely spoken languages of Africa, Asia, and South America, one endangered language of Central America, and three critically endangered languages of North America.
Outcome: The proposed architecture outperforms four of the most popular ASR toolkits for eleven languages with limited training resources.
Towards Language-Agnostic STIPA: Universal Phonetic Transcription to Support Language Documentation at Scale (2025.emnlp-main)

Copied to clipboard

Challenge: Existing ASR systems focus on orthographic output for high-resource languages, but STIPA can be used as a language-agnostic interface for documenting under-resourced and unwritten languages.
Approach: They propose to use the International Phonetic Alphabet (STIPA) to generate phonetic transcriptions using a language-agnostic interface.
Outcome: The proposed model reduces phonetic error rates even in low-resource settings and can be used for documenting under-resourced and unwritten languages.
ASR for Documenting Acutely Under-Resourced Indigenous Languages (L18-1)

Copied to clipboard

Challenge: Automatic speech recognition (ASR) has not been widely explored as a tool for documenting endangered languages.
Approach: They propose to use automatic speech recognition (ASR) to bootstrap new data to improve the acoustic model.
Outcome: The proposed system improves the model for a polysynthetic language with few audio and text resources.
How Important is a Language Model for Low-resource ASR? (2024.findings-acl)

Copied to clipboard

Challenge: Using an n-gram language model in ASR may seem obvious, but its absence in most implementations suggests otherwise.
Approach: They examine whether using an n-gram language model in ASR can improve accuracy in low-resource languages.
Outcome: The proposed model is absent in most implementations, but it does improve accuracy in English and Mandarin.
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)

Copied to clipboard

Challenge: Using acoustic data, we develop automatic speech recognition systems for three low resource languages.
Approach: They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework.
Outcome: The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut.
Can Large Language Models Translate Unseen Languages in Underrepresented Scripts? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive performance in machine translation, but struggle with unseen low-resource languages.
Approach: They propose a benchmark to evaluate translation for Mongolian and Yi using linguistic resources.
Outcome: The proposed model can translate Mongolian (in traditional script) and Yi with the help of linguistic resources, but is limited in its ability to handle these languages effectively.
Part-of-speech Tagging for Extremely Low-resource Indian Languages (2024.findings-acl)

Copied to clipboard

Challenge: Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation.
Approach: They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi.
Outcome: The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge.
High-quality Data-to-Text Generation for Severely Under-Resourced Languages with Out-of-the-box Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Pretrained large language models (LLMs) can bridge the performance gap for under-resourced languages by substantial margins, as measured by both automatic and human evaluations.
Approach: They propose to use pretrained large language models to bridge this gap by automating and evaluating data-to-text generation in under-resourced languages.
Outcome: The proposed model can set the state of the art for under-resourced languages by substantial margins, as measured by both automatic and human evaluations.
Thesis Proposal: Development of End-to-End Speech Translation Models for Indian Languages (2026.eacl-srw)

Copied to clipboard

Challenge: Existing approaches to speech-to-speech translation rely on cascaded pipelines . current approaches rely only on text representations, but they suffer from errors and latency . a new direct speech translation framework is proposed to bridge linguistic gaps .
Approach: They propose a sequence-to-sequence direct speech translation framework that can translate speech from one Indian language to another without relying on intermediate text representations.
Outcome: The proposed framework can translate speech from one Indian language to another without relying on intermediate text representations.
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages (2025.acl-long)

Copied to clipboard

Challenge: Existing datasets that cover only a fraction of Indian languages lack the breadth needed to generalize beyond curated benchmarks.
Approach: They propose to build the largest speech translation dataset for Indian languages . they use a three-step methodology to gather data and train a model that performs better .
Outcome: The proposed model improves on existing models and is open-source with permissive licenses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations