Language and Speech Technology for Central Kurdish Varieties (2024.lrec-main)

Copied to clipboard

Challenge: a recent study focused on the Kurdish language, a less-resourced Indo-European language spoken by over 30 million speakers.
Approach: They propose to develop resources for language and speech technology for Kurdish . they report the performance of machine translation, automatic speech recognition and language identification .
Outcome: The proposed model is based on transcribing movies and TV series as an alternative to fieldwork.

Similar Papers

Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)

Copied to clipboard

Challenge: Using acoustic data, we develop automatic speech recognition systems for three low resource languages.
Approach: They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework.
Outcome: The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
A description and demonstration of SAFAR framework (2021.eacl-demos)

Copied to clipboard

Challenge: Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language .
Approach: They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework"
Outcome: The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect.
What Do Indonesians Really Need from Language Technology? A Nationwide Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Despite efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers.
Approach: They conduct a nationwide survey to assess the actual needs of native Indonesian speakers.
Outcome: The findings indicate that addressing language barriers is the most critical priority . concerns around privacy, bias, and the use of public data highlight the need for greater transparency and clear communication to support broader AI adoption.
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)

Copied to clipboard

Challenge: Various corpora of various sizes and representing different genres, have been created for various Arabic dialects.
Approach: They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files.
Outcome: The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP .
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: despite recent advances in speech processing, the majority of world languages and dialects remain uncovered.
Approach: They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset .
Outcome: The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni.
Habibi - a multi Dialect multi National Arabic Song Lyrics Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Unlike western music, Arabic songs are poorly classified and the majority of the songs available online are classified under Modern Arabic Pop genre or what is now known as Franco-Arabic .
Approach: They introduce Habibi the first Arabic Song Lyrics corpus for singers from 18 different Arabic countries.
Outcome: The proposed corpus contains more than 30,000 Arabic song lyrics in 6 Arabic dialects for singers from 18 different arab countries.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Burmese Speech Corpus, Finite-State Text Normalization and Pronunciation Grammars with an Application to Text-to-Speech (2020.lrec-1)

Copied to clipboard

Challenge: Using crowd-sourced speech corpus and finite-state transducer grammars, we build a text-to-speech system for Burmese, a tonal Southeast Asian language from the Sino-Tibetan family.
Approach: They propose an open-source crowd-sourced multi-speaker speech corpus and finite-state grammars for performing grapheme-to-phoneme conversion for Burmese.
Outcome: The proposed system performs well for Burmese in a low-resource setting.
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic (2025.findings-acl)

Copied to clipboard

Challenge: Dialectal Arabic (DA) varieties are under-served by language technologies, particularly large language models (LLMs).
Approach: They propose a framework that comprehensively assesses LLMs’ DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia.
Outcome: The proposed framework assesses LLMs’ DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations