| Challenge: | a recent study focused on the Kurdish language, a less-resourced Indo-European language spoken by over 30 million speakers. |
| Approach: | They propose to develop resources for language and speech technology for Kurdish . they report the performance of machine translation, automatic speech recognition and language identification . |
| Outcome: | The proposed model is based on transcribing movies and TV series as an alternative to fieldwork. |
Similar Papers
Progress in Multilingual Speech Recognition for Low Resource Languages Kurmanji Kurdish, Cree and Inuktut (2022.lrec-1)
Copied to clipboard
| Challenge: | Using acoustic data, we develop automatic speech recognition systems for three low resource languages. |
| Approach: | They develop automatic speech recognition systems for three low resource languages using acoustic training data from 12 different languages in the hybrid DNN/HMM framework. |
| Outcome: | The proposed models are for three low resource languages: Kurmanji Kurdish, Cree and Inuktut. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
A description and demonstration of SAFAR framework (2021.eacl-demos)
Copied to clipboard
Karim Bouzoubaa, Younes Jaafar, Driss Namly, Ridouane Tachicart, Rachida Tajmout, Hakima Khamar, Hamid Jaafar, Lhoussain Aouragh, Abdellah Yousfi
| Challenge: | Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language . |
| Approach: | They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework" |
| Outcome: | The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect. |
What Do Indonesians Really Need from Language Technology? A Nationwide Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite efforts to develop NLP for Indonesia’s 700+ local languages, progress remains costly due to the need for direct engagement with native speakers. |
| Approach: | They conduct a nationwide survey to assess the actual needs of native Indonesian speakers. |
| Outcome: | The findings indicate that addressing language barriers is the most critical priority . concerns around privacy, bias, and the use of public data highlight the need for greater transparency and clear communication to support broader AI adoption. |
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)
Copied to clipboard
| Challenge: | Various corpora of various sizes and representing different genres, have been created for various Arabic dialects. |
| Approach: | They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files. |
| Outcome: | The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP . |
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)
Copied to clipboard
Bashar Talafha, Karima Kadaoui, Samar Magdy, Mariem Habiboullah, Chafei Chafei, Ahmed El-Shangiti, Hiba Zayed, Mohamedou Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed
| Challenge: | despite recent advances in speech processing, the majority of world languages and dialects remain uncovered. |
| Approach: | They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset . |
| Outcome: | The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. |
Habibi - a multi Dialect multi National Arabic Song Lyrics Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Unlike western music, Arabic songs are poorly classified and the majority of the songs available online are classified under Modern Arabic Pop genre or what is now known as Franco-Arabic . |
| Approach: | They introduce Habibi the first Arabic Song Lyrics corpus for singers from 18 different Arabic countries. |
| Outcome: | The proposed corpus contains more than 30,000 Arabic song lyrics in 6 Arabic dialects for singers from 18 different arab countries. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Burmese Speech Corpus, Finite-State Text Normalization and Pronunciation Grammars with an Application to Text-to-Speech (2020.lrec-1)
Copied to clipboard
Yin May Oo, Theeraphol Wattanavekin, Chenfang Li, Pasindu De Silva, Supheakmungkol Sarin, Knot Pipatsrisawat, Martin Jansche, Oddur Kjartansson, Alexander Gutkin
| Challenge: | Using crowd-sourced speech corpus and finite-state transducer grammars, we build a text-to-speech system for Burmese, a tonal Southeast Asian language from the Sino-Tibetan family. |
| Approach: | They propose an open-source crowd-sourced multi-speaker speech corpus and finite-state grammars for performing grapheme-to-phoneme conversion for Burmese. |
| Outcome: | The proposed system performs well for Burmese in a low-resource setting. |
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic (2025.findings-acl)
Copied to clipboard
| Challenge: | Dialectal Arabic (DA) varieties are under-served by language technologies, particularly large language models (LLMs). |
| Approach: | They propose a framework that comprehensively assesses LLMs’ DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia. |
| Outcome: | The proposed framework assesses LLMs’ DA modeling capabilities across four dimensions: fidelity, understanding, quality, and diglossia. |