Annotating Attribution Relations in Arabic (L18-1)

Copied to clipboard

Challenge: Current studies focus on using lexical terms in long texts to verify author identity.
Approach: They propose to annotate attributed arguments to the source in Arabic news with required syntactical and semantic features with required features.
Outcome: The proposed method is applied to Arabic news and is compared with existing tools and methods.

Similar Papers

BERT-based Classical Arabic Poetry Authorship Attribution (2025.coling-main)

Copied to clipboard

Challenge: AA in Arabic poetry has been a significant issue since the 9th century due to the loss of pre-Islamic poetry and the misattribution of post-Islamical works to earlier poets.
Approach: They propose a computational approach to authorship attribution in Arabic poetry using the entire Classical Arabic Poetry corpus for the first time.
Outcome: The proposed model achieves F1 scores ranging from 0.97 to 1.0 and was applied to four pre-Islamic misattribution cases.
Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)

Copied to clipboard

Challenge: a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently.
Approach: They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement.
Outcome: The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite.
AdabNER: Arabic Digital Archive Books with Nested Entity Recognition (2026.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a subtask of information extraction that classifies entities into predefined categories like person names.
Approach: They propose a large-scale nested Arabic Named Entity Recognition dataset . they fine-tuned five pre-trained Arabic BERT encoders in two settings .
Outcome: The first large-scale nested NER dataset for Arabic literary texts is published online . the dataset yields 78,530 entity mentions, 18.96% of which are nestated .
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)

Copied to clipboard

Challenge: Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA).
Approach: They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use.
Outcome: The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety.
Arabic Natural Language Processing (2022.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial provides background information for system developers and researchers working with Arabic in its various forms.
Approach: This tutorial provides the necessary background information for working with Arabic in its various forms.
Outcome: This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing.
Automatic Identification of Maghreb Dialects Using a Dictionary-Based Approach (L18-1)

Copied to clipboard

Challenge: Automatic identification of Arabic dialects in texts is difficult, especially for Maghreb languages and when they are written in Arabic or Latin characters (Arabizi).
Approach: They propose a dictionary-based approach to detect Arabic dialects in texts . they focus on transliteration of Arabicizi into Latin script and code-switching .
Outcome: The proposed approach shows that it is possible to detect dialects in Arabic and Latin scripts.
Masader: Metadata Sourcing for Arabic Text and Speech Data Resources (2022.lrec-1)

Copied to clipboard

Challenge: Currently, there is no online catalogue for Arabic datasets with annotated attributes . this paper aims to identify the publicly available Arabic dataset and provide a catalogue of them to researchers.
Approach: They propose to create the largest public catalogue for Arabic NLP datasets with 25 attributes and a metadata annotation strategy that could be extended to other languages.
Outcome: The proposed approach could be extended to other languages and regions.
An Attribution Relations Corpus for Political News (L18-1)

Copied to clipboard

Challenge: Existing resources for recognizing attributions in context are limited in size and completeness.
Approach: They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction .
Outcome: The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date.
Authorship Attribution in Multilingual Machine-Generated Texts (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have reached human-like fluency and coherence, but distinguishing machine-generated text from human-written content becomes increasingly difficult.
Approach: They propose a problem of multilingual authorship attribution (AA) that involves attributing texts to human or multiple LLM generators across diverse languages.
Outcome: The proposed method can be adapted to multilingual settings, but still has significant limitations and challenges.
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)

Copied to clipboard

Challenge: CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses)
Approach: They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic.
Outcome: The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses)

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations