| Challenge: | ZAEBUC is an annotated Arabic-English bilingual writer corpus . it is a corpus of short essays written by first-year university students . |
| Approach: | They propose to use a standard Arabic-English bilingual writer corpus to match comparable texts written by the same writer on different occasions. |
| Outcome: | The ZAEBUC corpus is an annotated Arabic-English bilingual writer corpus by first-year university students at Zayed University in the United Arab Emirates. |
Similar Papers
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | a corpus of multilingual Arabic-English speech is presented in a new paper . a major bottleneck is the lack of data needed for training NLP models . |
| Approach: | They propose a multilingual multidialectal Arabic-English speech corpus with a set of guidelines for automatic speech recognition. |
| Outcome: | The proposed corpus includes two languages with Arabic and English spoken in multiple variants and Arabic and Arabic with various accents. |
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)
Copied to clipboard
| Challenge: | Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA). |
| Approach: | They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use. |
| Outcome: | The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety. |
Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)
Copied to clipboard
Najet Hadj Mohamed, Cherifa Ben Khelil, Agata Savary, Iskandar Keskes, Jean-Yves Antoine, Lamia Hadrich-Belguith
| Challenge: | a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently. |
| Approach: | They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement. |
| Outcome: | The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite. |
EMAD: A Bridge Tagset for Unifying Arabic POS Annotations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing tagsets for Arabic are difficult to combine due to the diversity of their features. |
| Approach: | They propose an Arabic Extended Morphological Analysis and Disambiguation Tagset which facilitates conversion and unification of Arabic tagsets. |
| Outcome: | The proposed tagset facilitates conversion and unification of different tagsetes used to annotate Arabic corpora. |
QCAW 1.0: Building a Qatari Corpus of Student Argumentative Writing (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have highlighted the importance of and need to create learner corpora. |
| Approach: | They propose to create a Qatari corpus of argumentative writing (QCAW) the corpus contains 200,000 tokens of argumentation written by Qatari university students . |
| Outcome: | The QCAW contains 195 essays written by 195 students, 159 females and 36 males. |
The SAMER Arabic Text Simplification Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Our corpus includes 159K words selected from 15 publicly available Arabic fiction novels . text simplification aims to reduce the complexity of a text while maintaining the overall grammaticality and core content. |
| Approach: | They propose to annotate Arabic parallel corpus for text simplification targeting school-aged learners. |
| Outcome: | The SAMER Corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels. |
A description and demonstration of SAFAR framework (2021.eacl-demos)
Copied to clipboard
Karim Bouzoubaa, Younes Jaafar, Driss Namly, Ridouane Tachicart, Rachida Tajmout, Hakima Khamar, Hamid Jaafar, Lhoussain Aouragh, Abdellah Yousfi
| Challenge: | Existing NLP infrastructures are naming them "toolkit", "platform" and "framework" authors present a monolingual framework dedicated to Arabic language . |
| Approach: | They propose a monolingual framework dedicated to Arabic language . they propose namings for existing infrastructures: "toolkit", "platform" and "framework" |
| Outcome: | The proposed framework is dedicated to Arabic language, especially the modern standard Arabic and Moroccan dialect. |
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)
Copied to clipboard
| Challenge: | Various corpora of various sizes and representing different genres, have been created for various Arabic dialects. |
| Approach: | They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files. |
| Outcome: | The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Code-switching is a phenomenon commonly observed in the Arabicspeaking world . there is still a huge gap in the available resources and NLP applications . |
| Approach: | They propose a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags. |
| Outcome: | The proposed corpus of Egyptian- Arabic code-switch speech data is fully tokenized, lemmatized and annotated for part-of-speech tags. |