Challenge: Part-of-speech (POS) tagging is one of the most important building blocks in many natural language processing (NLP) applications.
Approach: They propose to use a POS tagger for Arabic Gulf dialect to improve POS tagging accuracy.
Outcome: The proposed POS tagger improves POS tagging accuracy for the Arabic Gulf dialect from 75% accuracy to 91% accuracy using a bi-LSTM labeler.

Similar Papers

Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)

Copied to clipboard

Challenge: Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability.
Approach: They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect.
Outcome: The proposed model can tag four different dialects with an average accuracy of 89.3%.
Active Learning for Multidialectal Arabic POS Tagging (2025.findings-emnlp)

Copied to clipboard

Challenge: Multidialectal Arabic POS tagging is challenging due to the morphological richness and high variability among dialects.
Approach: They propose an active learning approach for multidialectal Arabic POS tagging . they annotate approximately 15,000 tokens, reducing the annotation requirement by about 2,000 tokens .
Outcome: The proposed approach achieves 97.6% accuracy on the Emirati corpus.
Classification of Closely Related Sub-dialects of Arabic Using Support-Vector Machines (L18-1)

Copied to clipboard

Challenge: Existing studies on dialect identification have focused on binary classifications between colloquial Arabic and dialectal Egyptian .
Approach: They propose to use an n-gram based SVM to classify on a fine-grained sub-dialectal level and compare it to methods used in dialect classification such as vocabulary pruning.
Outcome: The proposed method is compared to methods used in dialect classification such as vocabulary pruning of shared items across dialects.
Highly Effective Arabic Diacritization using Sequence to Sequence Modeling (N19-1)

Copied to clipboard

Challenge: Arabic text is written without short vowels (or diacritics) their presence is essential for properly verbalizing Arabic .
Approach: They propose a character-level sequence-to-sequence deep learning model that recovers both types of diacritics without the use of explicit feature engineering.
Outcome: The proposed model outperforms all previous state-of-the-art models on overlapping windows of words . it achieves a word error rate (WER) of 4.49% compared to the state- of-the art systems .
Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained morphosyntactic tagging models outperform existing systems in Modern Standard Arabic and all the Arabic dialects studied.
Approach: They present results on morphosyntactic tagging across different varieties of Arabic using pre-trained transformer language models.
Outcome: The proposed models outperform existing systems in Modern Standard Arabic, 2.8% in Gulf, 1.6% in Egyptian, and 8.3% in Levantine.
A System for Diacritizing Four Varieties of Arabic (D19-3)

Copied to clipboard

Challenge: Short vowels, aka diacritics, are omitted when writing different varieties of Arabic . diacritization is essential for language learning and text-to-speech applications .
Approach: They propose a system for recovering diacritics in Arabic without short vowels . they use a character-based sequence-to-sequence deep learning model .
Outcome: The proposed system beats all previous SOTA systems for Arabic varieties . it uses a character-based sequence-to-sequence deep learning model .
Part-of-speech Tagging for Extremely Low-resource Indian Languages (2024.findings-acl)

Copied to clipboard

Challenge: Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation.
Approach: They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi.
Outcome: The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge.
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: despite recent advances in speech processing, the majority of world languages and dialects remain uncovered.
Approach: They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset .
Outcome: The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni.
Dialectal Coverage And Generalization in Arabic Speech Recognition (2025.acl-long)

Copied to clipboard

Challenge: Existing ASR systems cover the modern standard Arabic variety but fail to cover the multitude of spoken variants.
Approach: They propose a suite of automatic speech recognition models optimized to recognize multiple variants of spoken Arabic.
Outcome: The proposed models show coverage and performance gains compared to prior models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations