| Challenge: | Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications. |
| Approach: | They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy. |
| Outcome: | The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler. |
Similar Papers
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)
Copied to clipboard
Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali, Mohamed Eldesouki, Younes Samih, Randah Alharbi, Mohammed Attia, Walid Magdy, Laura Kallmeyer
| Challenge: | Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability. |
| Approach: | They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect. |
| Outcome: | The proposed model can tag four different dialects with an average accuracy of 89.3%. |
Active Learning for Multidialectal Arabic POS Tagging (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Multidialectal Arabic POS tagging is challenging due to the morphological richness and high variability among dialects. |
| Approach: | They propose an active learning approach for multidialectal Arabic POS tagging . they annotate approximately 15,000 tokens, reducing the annotation requirement by about 2,000 tokens . |
| Outcome: | The proposed approach achieves 97.6% accuracy on the Emirati corpus. |
Classification of Closely Related Sub-dialects of Arabic Using Support-Vector Machines (L18-1)
Copied to clipboard
| Challenge: | Existing studies on dialect identification have focused on binary classifications between colloquial Arabic and dialectal Egyptian . |
| Approach: | They propose to use an n-gram based SVM to classify on a fine-grained sub-dialectal level and compare it to methods used in dialect classification such as vocabulary pruning. |
| Outcome: | The proposed method is compared to methods used in dialect classification such as vocabulary pruning of shared items across dialects. |
MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African languages (2023.acl-long)
Copied to clipboard
Cheikh M. Bamba Dione, David Ifeoluwa Adelani, Peter Nabende, Jesujoba Alabi, Thapelo Sindane, Happy Buzaaba, Shamsuddeen Hassan Muhammad, Chris Chinenye Emezue, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jonathan Mukiibi, Blessing Sibanda, Bonaventure F. P. Dossou, Andiswa Bukula, Rooweither Mabuya, Allahsera Auguste Tapo, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Fatoumata Ouoba Kabore, Amelia Taylor, Godson Kalipe, Tebogo Macucwa, Vukosi Marivate, Tajuddeen Gwadabe, Mboning Tchiaze Elvis, Ikechukwu Onyenwe, Gratien Atindogbe, Tolulope Adelani, Idris Akinade, Olanrewaju Samuel, Marien Nahimana, Théogène Musabeyezu, Emile Niyomutabazi, Ester Chimhenga, Kudzai Gotosa, Patrick Mizha, Apelete Agbolo, Seydou Traore, Chinedu Uchechukwu, Aliyu Yusuf, Muhammad Abdullahi, Dietrich Klakow
| Challenge: | POS tagging is one of the fundamental steps for many natural language processing (NLP) applications. |
| Approach: | They present AfricaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. |
| Outcome: | The proposed model improves POS tagging performance in unseen languages. |
Highly Effective Arabic Diacritization using Sequence to Sequence Modeling (N19-1)
Copied to clipboard
| Challenge: | Arabic text is written without short vowels (or diacritics) their presence is essential for properly verbalizing Arabic . |
| Approach: | They propose a character-level sequence-to-sequence deep learning model that recovers both types of diacritics without the use of explicit feature engineering. |
| Outcome: | The proposed model outperforms all previous state-of-the-art models on overlapping windows of words . it achieves a word error rate (WER) of 4.49% compared to the state- of-the art systems . |
Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects (2022.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained morphosyntactic tagging models outperform existing systems in Modern Standard Arabic and all the Arabic dialects studied. |
| Approach: | They present results on morphosyntactic tagging across different varieties of Arabic using pre-trained transformer language models. |
| Outcome: | The proposed models outperform existing systems in Modern Standard Arabic, 2.8% in Gulf, 1.6% in Egyptian, and 8.3% in Levantine. |
A System for Diacritizing Four Varieties of Arabic (D19-3)
Copied to clipboard
| Challenge: | Short vowels, aka diacritics, are omitted when writing different varieties of Arabic . diacritization is essential for language learning and text-to-speech applications . |
| Approach: | They propose a system for recovering diacritics in Arabic without short vowels . they use a character-based sequence-to-sequence deep learning model . |
| Outcome: | The proposed system beats all previous SOTA systems for Arabic varieties . it uses a character-based sequence-to-sequence deep learning model . |
Part-of-speech Tagging for Extremely Low-resource Indian Languages (2024.findings-acl)
Copied to clipboard
| Challenge: | Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation. |
| Approach: | They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi. |
| Outcome: | The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge. |
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)
Copied to clipboard
Bashar Talafha, Karima Kadaoui, Samar Magdy, Mariem Habiboullah, Chafei Chafei, Ahmed El-Shangiti, Hiba Zayed, Mohamedou Tourad, Rahaf Alhamouri, Rwaa Assi, Aisha Alraeesi, Hour Mohamed, Fakhraddin Alwajih, Abdelrahman Mohamed, Abdellah El Mekki, El Moatez Billah Nagoudi, Benelhadj Saadia, Hamzah Alsayadi, Walid Al-Dhabyani, Sara Shatnawi, Yasir Ech-chammakhy, Amal Makouar, Yousra Berrachedi, Mustafa Jarrar, Shady Shehata, Ismail Berrada, Muhammad Abdul-Mageed
| Challenge: | despite recent advances in speech processing, the majority of world languages and dialects remain uncovered. |
| Approach: | They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset . |
| Outcome: | The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni. |
Dialectal Coverage And Generalization in Arabic Speech Recognition (2025.acl-long)
Copied to clipboard
| Challenge: | Existing ASR systems cover the modern standard Arabic variety but fail to cover the multitude of spoken variants. |
| Approach: | They propose a suite of automatic speech recognition models optimized to recognize multiple variants of spoken Arabic. |
| Outcome: | The proposed models show coverage and performance gains compared to prior models. |