| Challenge: | The processing of the Arabic language is a complex field of research due to the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need to be processed while taking into account their unique characteristics. |
| Approach: | They propose to revise the Palestinian morphologically annotated corpus and a Lebanese corpus to bridge nuanced linguistic gaps between the two highly mutually intelligible dialects. |
| Outcome: | The revised corpus can be used as a more general Levantine corpus. |
Similar Papers
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)
Copied to clipboard
| Challenge: | Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA). |
| Approach: | They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use. |
| Outcome: | The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety. |
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)
Copied to clipboard
| Challenge: | Various corpora of various sizes and representing different genres, have been created for various Arabic dialects. |
| Approach: | They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files. |
| Outcome: | The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP . |
EMAD: A Bridge Tagset for Unifying Arabic POS Annotations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing tagsets for Arabic are difficult to combine due to the diversity of their features. |
| Approach: | They propose an Arabic Extended Morphological Analysis and Disambiguation Tagset which facilitates conversion and unification of Arabic tagsets. |
| Outcome: | The proposed tagset facilitates conversion and unification of different tagsetes used to annotate Arabic corpora. |
Shami: A Corpus of Levantine Arabic Dialects (L18-1)
Copied to clipboard
| Challenge: | Modern Standard Arabic is the official written language used in education and media . however, the spoken language varies widely across the Arab world . |
| Approach: | They construct a levantine dialect corpus covering data from four dialects spoken in four countries . they describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools. |
| Outcome: | The proposed corpus is larger than existing corpora in terms of size, words and vocabularies. |
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)
Copied to clipboard
| Challenge: | CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses) |
| Approach: | They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic. |
| Outcome: | The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses) |
Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging (2020.acl-main)
Copied to clipboard
| Challenge: | a word can have multiple interpretations and is one of many inflected forms of the same concept or lemma. |
| Approach: | They propose to model morphological features jointly, whether lexicalized or non-lexicalised . their results are compared to Arabic and Egyptian Arabic . |
| Outcome: | The proposed model achieves 20% relative error reduction in Arabic and 11% in Egyptian Arabic. |
TArC: Tunisian Arabish Corpus, First complete release (2022.lrec-1)
Copied to clipboard
| Challenge: | a project focused on Tunisian Arabic encoded in Arabizi is a hybrid approach to linguistics and linguistic research . Arabic dialects are notoriously under-resourced linguistic systems . |
| Approach: | They propose to use Arabic script as a linguistic corpus and a neural network architecture to annotate the latter with various levels of linguistic information. |
| Outcome: | The proposed approach is hybrid and combines linguistic and linguistic tools . the proposed approach produces in cascade different levels of annotation . |
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment (2025.findings-acl)
Copied to clipboard
| Challenge: | Texts above a student's readability level can lead to disengagement and disengagement . Developing readability models is crucial for improving literacy, language learning, and academic performance. |
| Approach: | They introduce the Balanced Arabic Readability Evaluation Corpus (BAREC) a large-scale, fine-grained dataset for Arabic readability assessment. |
| Outcome: | The proposed model outperforms existing methods in Arabic readability assessment. |
A Leveled Reading Corpus of Modern Standard Arabic (L18-1)
Copied to clipboard
| Challenge: | Using a reading corpus in Modern Standard Arabic, we explore the lexical coverage of textbooks and unabridged works of fiction. |
| Approach: | They propose to use textbooks from the United Arab Emirates curriculum and a reading corpus in Modern Standard Arabic to enrich the sparse collection of resources available for educational applications. |
| Outcome: | The corpus spans all 12 grades and contains 129 unabridged works of fiction spanning grades 1-12 . lexical coverage is compared to other genres, and the results show that the two sub-corpora are similar to each other to measure their genres. |
Constructing a Bilingual Hadith Corpus Using a Segmentation Tool (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on Hadith have focused on the Quran, leaving it relatively unexplored. |
| Approach: | They propose to gather and construct a bilingual parallel corpus of Islamic Hadith using a custom segmentation tool that annotates the two Hadithe components with 92% accuracy. |
| Outcome: | The proposed method minimises the costs of language resource creation and produces consistent results independently from previous knowledge and experiences that usually influence human annotators. |