Annotating Verbal Multiword Expressions in Arabic: Assessing the Validity of a Multilingual Annotation Procedure (2022.lrec-1)
Copied to clipboard
Najet Hadj Mohamed, Cherifa Ben Khelil, Agata Savary, Iskandar Keskes, Jean-Yves Antoine, Lamia Hadrich-Belguith
| Challenge: | a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT were selected and annotated by two Arabic native speakers independently. |
| Approach: | They propose to use Arabic as an annotation framework to extend PARSEME to modern standard Arabic by measuring inter-annotator agreement. |
| Outcome: | The proposed framework is based on a subset of 1,062 sentences from the Prague Arabic Dependency Treebank PADT and is already exceeding the smallest corpus of the PARSEME suite. |
Similar Papers
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)
Copied to clipboard
| Challenge: | Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA). |
| Approach: | They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use. |
| Outcome: | The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety. |
ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | ZAEBUC is an annotated Arabic-English bilingual writer corpus . it is a corpus of short essays written by first-year university students . |
| Approach: | They propose to use a standard Arabic-English bilingual writer corpus to match comparable texts written by the same writer on different occasions. |
| Outcome: | The ZAEBUC corpus is an annotated Arabic-English bilingual writer corpus by first-year university students at Zayed University in the United Arab Emirates. |
Some Issues with Building a Multilingual Wordnet (2020.lrec-1)
Copied to clipboard
| Challenge: | Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalised synsets and lemmas, new parts of speech, and more. |
| Approach: | They propose to integrate a new open multilingual wordnet format that tests the extensions introduced by the new format and integrates a set of tools to ensure the integrity of the Collaborative Interlingual Index. |
| Outcome: | The proposed format integrates a set of tools that test the extensions while ensuring the integrity of the Collaborative Interlingual Index (CILI). |
The SAMER Arabic Text Simplification Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Our corpus includes 159K words selected from 15 publicly available Arabic fiction novels . text simplification aims to reduce the complexity of a text while maintaining the overall grammaticality and core content. |
| Approach: | They propose to annotate Arabic parallel corpus for text simplification targeting school-aged learners. |
| Outcome: | The SAMER Corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels. |
A Large-Scale Leveled Readability Lexicon for Standard Arabic (2020.lrec-1)
Copied to clipboard
| Challenge: | a large-scale leveled readability lexicon for Modern Standard Arabic is not available in many other languages. |
| Approach: | They propose a large-scale leveled readability lexicon for Arabic with 26,000 lemmas . they manually annotate a lexico from three different regions in the arab world . |
| Outcome: | The proposed lexicon is publicly available for Arabic readability tasks. |
Annotating Attribution Relations in Arabic (L18-1)
Copied to clipboard
| Challenge: | Current studies focus on using lexical terms in long texts to verify author identity. |
| Approach: | They propose to annotate attributed arguments to the source in Arabic news with required syntactical and semantic features with required features. |
| Outcome: | The proposed method is applied to Arabic news and is compared with existing tools and methods. |
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic (2024.findings-acl)
Copied to clipboard
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, Timothy Baldwin
| Challenge: | evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets. |
| Approach: | They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries . |
| Outcome: | The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions . |
Construction of Large-scale English Verbal Multiword Expression Annotated Corpus (L18-1)
Copied to clipboard
| Challenge: | In this paper, we focus on verbal MWEs, whose accurate recognition is challenging because they could be discontinuous. |
| Approach: | They conduct large-scale annotations of VMWEs on the Wall Street Journal portion of Ontonotes . they first construct a VMwe dictionary based on the english-language Wiktionary . |
| Outcome: | The proposed resource annotates 7,833 VMWE instances belonging to various categories . the authors hope the results will help to develop models for MWE recognition and dependency parsing . |
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)
Copied to clipboard
| Challenge: | Terms are notoriously difficult to identify, both automatically and manually. |
| Approach: | They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information . |
| Outcome: | The proposed method provides a tool for evaluation and rich source of information about terms. |
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)
Copied to clipboard
| Challenge: | CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses) |
| Approach: | They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic. |
| Outcome: | The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses) |