| Challenge: | Current studies focus on using lexical terms in long texts to verify author identity. |
| Approach: | They propose to annotate attributed arguments to the source in Arabic news with required syntactical and semantic features with required features. |
| Outcome: | The proposed method is applied to Arabic news and is compared with existing tools and methods. |
Similar Papers
BERT-based Classical Arabic Poetry Authorship Attribution (2025.coling-main)
Copied to clipboard
| Challenge: | AA in Arabic poetry has been a significant issue since the 9th century due to the loss of pre-Islamic poetry and the misattribution of post-Islamical works to earlier poets. |
| Approach: | They propose a computational approach to authorship attribution in Arabic poetry using the entire Classical Arabic Poetry corpus for the first time. |
| Outcome: | The proposed model achieves F1 scores ranging from 0.97 to 1.0 and was applied to four pre-Islamic misattribution cases. |
Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)
Copied to clipboard
Najet Hadj Mohamed, Cherifa Ben Khelil, Agata Savary, Iskandar Keskes, Jean-Yves Antoine, Lamia Hadrich-Belguith
| Challenge: | a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently. |
| Approach: | They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement. |
| Outcome: | The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite. |
AdabNER: Arabic Digital Archive Books with Nested Entity Recognition (2026.acl-long)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a subtask of information extraction that classifies entities into predefined categories like person names. |
| Approach: | They propose a large-scale nested Arabic Named Entity Recognition dataset . they fine-tuned five pre-trained Arabic BERT encoders in two settings . |
| Outcome: | The first large-scale nested NER dataset for Arabic literary texts is published online . the dataset yields 78,530 entity mentions, 18.96% of which are nestated . |
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)
Copied to clipboard
| Challenge: | Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA). |
| Approach: | They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use. |
| Outcome: | The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety. |
Arabic Natural Language Processing (2022.emnlp-tutorials)
Copied to clipboard
| Challenge: | This tutorial provides background information for system developers and researchers working with Arabic in its various forms. |
| Approach: | This tutorial provides the necessary background information for working with Arabic in its various forms. |
| Outcome: | This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing. |
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach (L18-1)
Copied to clipboard
| Challenge: | Automatic identification of Arabic dialects in texts is difficult, especially for Maghreb languages and when they are written in Arabic or Latin characters (Arabizi). |
| Approach: | They propose a dictionary-based approach to detect Arabic dialects in texts . they focus on transliteration of Arabicizi into Latin script and code-switching . |
| Outcome: | The proposed approach shows that it is possible to detect dialects in Arabic and Latin scripts. |
Masader: Metadata Sourcing for Arabic Text and Speech Data Resources (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, there is no online catalogue for Arabic datasets with annotated attributes . this paper aims to identify the publicly available Arabic dataset and provide a catalogue of them to researchers. |
| Approach: | They propose to create the largest public catalogue for Arabic NLP datasets with 25 attributes and a metadata annotation strategy that could be extended to other languages. |
| Outcome: | The proposed approach could be extended to other languages and regions. |
An Attribution Relations Corpus for Political News (L18-1)
Copied to clipboard
| Challenge: | Existing resources for recognizing attributions in context are limited in size and completeness. |
| Approach: | They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction . |
| Outcome: | The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date. |
Authorship Attribution in Multilingual Machine-Generated Texts (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have reached human-like fluency and coherence, but distinguishing machine-generated text from human-written content becomes increasingly difficult. |
| Approach: | They propose a problem of multilingual authorship attribution (AA) that involves attributing texts to human or multiple LLM generators across diverse languages. |
| Outcome: | The proposed method can be adapted to multilingual settings, but still has significant limitations and challenges. |
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)
Copied to clipboard
| Challenge: | CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses) |
| Approach: | They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic. |
| Outcome: | The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses) |