Papers by Nizar Habash
Copied to clipboard
| Challenge: | Recent work utilizing a mix of pattern-based and distributional approaches have yielded state-of-the-art results on two domain-specific English hypernym discovery tasks. |
| Approach: | They evaluate the contribution of pattern-based and distributional approaches to hybrid modeling by evaluating baseline models from each paradigm. |
| Outcome: | The proposed approach outperforms all non-hybrid approaches on two domain-specific English hypernym discovery tasks and outperformed other approaches. |
Copied to clipboard
| Challenge: | Diacritics can significantly influence language processing tasks in Arabic . their presence can increase subword fragmentation during tokenization, reducing performance . |
| Approach: | They analyze the impact of diacritics on tokenization and benchmark task performance across major Large Language Models. |
| Outcome: | The proposed model is robust to diacritics, but full diacritization leads to token fragmentation and degraded performance. |
Copied to clipboard
| Challenge: | Texts above a student's readability level can lead to disengagement and disengagement . Developing readability models is crucial for improving literacy, language learning, and academic performance. |
| Approach: | They introduce the Balanced Arabic Readability Evaluation Corpus (BAREC) a large-scale, fine-grained dataset for Arabic readability assessment. |
| Outcome: | The proposed model outperforms existing methods in Arabic readability assessment. |
Copied to clipboard
| Challenge: | Maltese is a Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. |
| Approach: | They investigate whether Arabic-language resources can support Maltese natural language processing . they introduce transliteration schemes and machine translation approaches to align Arabic text with Maltesen . |
| Outcome: | The proposed techniques can significantly improve Maltese natural language processing tasks. |
Copied to clipboard
| Challenge: | This tutorial provides background information for system developers and researchers working with Arabic in its various forms. |
| Approach: | This tutorial provides the necessary background information for working with Arabic in its various forms. |
| Outcome: | This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing. |
Copied to clipboard
| Challenge: | Recent work addresses this issue by modeling dialectness as a continuous variable . however, ALDi reduces complex variation to a single dimension . |
| Approach: | They propose a way to model Arabic dialectness as a continuous variable . they propose etymology-aware edit distance and a regression model to model AGS . |
| Outcome: | The proposed approach outperforms baselines on a multi-dialect benchmark. |
Copied to clipboard
| Challenge: | Code-switching (CS) is a problem in machine translation, but its performance is not investigated for CS settings. |
| Approach: | They propose to use morphological segmentation techniques for machine translation tasks . they compare morphology-based and frequency-based segmentation for MT tasks based on data size . |
| Outcome: | The proposed approach performs best in MT tasks but under-performs in other languages. |
Copied to clipboard
| Challenge: | a word can have multiple interpretations and is one of many inflected forms of the same concept or lemma. |
| Approach: | They propose to model morphological features jointly, whether lexicalized or non-lexicalised . their results are compared to Arabic and Egyptian Arabic . |
| Outcome: | The proposed model achieves 20% relative error reduction in Arabic and 11% in Egyptian Arabic. |
Copied to clipboard
| Challenge: | ZAEBUC is an annotated Arabic-English bilingual writer corpus . it is a corpus of short essays written by first-year university students . |
| Approach: | They propose to use a standard Arabic-English bilingual writer corpus to match comparable texts written by the same writer on different occasions. |
| Outcome: | The ZAEBUC corpus is an annotated Arabic-English bilingual writer corpus by first-year university students at Zayed University in the United Arab Emirates. |
Copied to clipboard
| Challenge: | Large language models generate fluent responses to user queries, but they are also susceptible to misuse in journalism, education, and academia. |
| Approach: | They propose a large-scale benchmark for machine-generated text detection that is a multi-generator, multi-domain, and multi-lingual corpus. |
| Outcome: | The proposed system can detect machine-generated text and pinpoint misuse . the proposed system is based on a large-scale benchmark dataset . |
Copied to clipboard
| Challenge: | Query Translator is a cross-lingual messaging app for the travel domain that automatically translates conversations . the application addresses common cross-linguistic communication issues such as translation accuracy, speed, privacy and personalization. |
| Approach: | They present a cross-lingual messaging app that automatically translates conversations while supporting keyword-to-sentence matching. |
| Outcome: | The proposed app translates conversations while supporting keyword-to-sentence matching. |
Copied to clipboard
| Challenge: | Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks. |
| Approach: | They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages. |
| Outcome: | The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks. |
Copied to clipboard
| Challenge: | Pre-trained morphosyntactic tagging models outperform existing systems in Modern Standard Arabic and all the Arabic dialects studied. |
| Approach: | They present results on morphosyntactic tagging across different varieties of Arabic using pre-trained transformer language models. |
| Outcome: | The proposed models outperform existing systems in Modern Standard Arabic, 2.8% in Gulf, 1.6% in Egyptian, and 8.3% in Levantine. |
Copied to clipboard
| Challenge: | a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs. |
| Approach: | They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus. |
| Outcome: | The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it . |
Copied to clipboard
| Challenge: | CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses) |
| Approach: | They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic. |
| Outcome: | The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses) |
Copied to clipboard
| Challenge: | Dialectal Arabic (DA) is problematically noisy and lacks a large corpus of non-noisy words. |
| Approach: | They propose to use word embedding tools to maximize the informative content leveraged in each training sentence and analyze methods for representing disparate dialects in one embeddable space. |
| Outcome: | The proposed methods improve performance on low and high frequency words while preserving accuracy on low frequency forms. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have brought an unprecedented surge in machine-generated text (MGT) societal implications are posed by their potential misuse and lack of training data. |
| Approach: | They propose a benchmark to detect machine-generated text in multiple languages . they use multi-domain and multi-generator corpus to identify which model generated the text . |
| Outcome: | The proposed benchmark compares a multilingual, multi-domain and multi-generator corpus of MGTs with human-generated content. |
Copied to clipboard
| Challenge: | Gender bias in natural language processing (NLP) applications has been receiving increasing attention, largely due to the lack of datasets and resources. |
| Approach: | They propose a corpus for gender identification and rewriting in contexts involving one or two target users with independent grammatical gender preferences. |
| Outcome: | The proposed corpus expands on Habash et al.'s Arabic Parallel Gender Corpus (APGC) by adding second person targets and increasing the total number of sentences over 6.5 times, reaching over 590K words. |
Copied to clipboard
| Challenge: | Camel Morph MSA is the largest open-source Modern Standard Arabic morphological analyzer and generator. |
| Approach: | They present Camel Morph MSA, the largest open-source Arabic morphological analyzer and generator. |
| Outcome: | The analysis can produce 1.45B analyses and 535M unique diacritizations, almost an order of magnitude larger than SAMA on a 10B word corpus. |
Copied to clipboard
| Challenge: | Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation. |
| Approach: | They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project. |
| Outcome: | The proposed resources are the first of their kind in terms of their coverage and fine granularity. |
Copied to clipboard
| Challenge: | Various corpora of various sizes and representing different genres, have been created for various Arabic dialects. |
| Approach: | They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files. |
| Outcome: | The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP . |
Copied to clipboard
| Challenge: | a corpus of multilingual Arabic-English speech is presented in a new paper . a major bottleneck is the lack of data needed for training NLP models . |
| Approach: | They propose a multilingual multidialectal Arabic-English speech corpus with a set of guidelines for automatic speech recognition. |
| Outcome: | The proposed corpus includes two languages with Arabic and English spoken in multiple variants and Arabic and Arabic with various accents. |
Copied to clipboard
| Challenge: | Existing systems that embed and amplify gender bias can still exhibit and exacerbate this problem. |
| Approach: | They propose a multi-step system that combines the positive aspects of rule-based and neural rewriting models to provide personalized outputs based on the users’ grammatical gender preferences. |
| Outcome: | The proposed system achieves 88.42 M2 F0.5 on a blind test set and improves over previous work on the first-person-only version of this task by 3.05 absolute increase in M2F0.5. |
Copied to clipboard
| Challenge: | Multilingual models exhibit impressive cross-lingual transfer capabilities on unseen languages, but performance is impacted when there is a script disparity with the languages used in the model’s pre-training data. |
| Approach: | They propose a novel method to align a resource-rich language's script with a target language and train a classifier that can make informed decisions regarding the appropriate processing of each token. |
| Outcome: | The proposed model can be used to transfer a language's scripts across multiple languages, but it is suboptimal for mixed languages, where only a subset benefits while the rest is impeded. |
Copied to clipboard
| Challenge: | Existing tools for lemmatization in morphologically rich languages with ambiguous orthography face inconsistent standards and limited genre coverage. |
| Approach: | They propose two new approaches that frame lemmatization as classification into a Lemma-POS-Gloss tagset, leveraging machine translation and semantic clustering. |
| Outcome: | The proposed models perform better than existing models and are more interpretable, the authors show. |
Copied to clipboard
| Challenge: | a Google Docs add-on for automatic Arabic word-level readability visualization is available for free. |
| Approach: | They propose a Google Docs add-on for automatic Arabic word-level readability visualization. |
| Outcome: | The proposed add-on can be used to assess the reading difficulty of a text and identify difficult words as part of manual text simplification. |
Copied to clipboard
| Challenge: | Prior studies have shown that distinguishing text generated by Large Language Models from human-written text is challenging for humans and often no better than random guessing. |
| Approach: | They conduct extensive case study to determine the upper bound of human detection accuracy. |
| Outcome: | The findings challenge previous conclusions on human detection accuracy across languages and domains. |
Copied to clipboard
| Challenge: | Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language . |
| Approach: | They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research . |
| Outcome: | This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research . |
Copied to clipboard
| Challenge: | a large number of machine-generated texts are often hard to distinguish between human-written and machine-generated text . this raises concerns about potential misuse, especially within educational and academic domains . |
| Approach: | They propose a system that can detect whether a text is human-written or machine-generated . they use a fine-grained classification schema to identify the use of machine-generated text . |
| Outcome: | The proposed system can distinguish between human-written and machine-generated text . it can detect attempts to obfuscate the fact that a text was machine- generated . |
Copied to clipboard
| Challenge: | Camelira is a web-based Arabic multi-dialect morphological disambiguation tool that covers modern standard Arabic, Egyptian, Gulf, and Levantine. |
| Approach: | They propose a web-based Arabic multi-dialect morphological disambiguation tool that covers modern standard Arabic, Egyptian, Gulf, and Levantine. |
| Outcome: | The proposed tool covers modern standard Arabic, Egyptian, Gulf, and Levantine . it also provides an option to automatically choose an appropriate disambiguator based on the prediction of a dialect identification component. |
Copied to clipboard
| Challenge: | Existing CDS corpora in English are limited due to their limited size or lack of linguistic annotation. |
| Approach: | They propose to map Egyptian Arabic IPA tokens to Arabic script and add core part-of-speech tags and lemmas aligned with existing Arabic morphological resources. |
| Outcome: | The proposed version of the Egyptian Arabic CHILDES corpus is available online. |
Copied to clipboard
| Challenge: | Noisy content is non-canonical in nature, with lexical, orthographic, and phonetic variations. |
| Approach: | They propose a neural morphological tagging and disambiguation model for Egyptian Arabic with various extensions to handle noisy content. |
| Outcome: | The proposed model achieves about 5% relative error reduction over a state-of-the-art baseline for Egyptian Arabic. |
Copied to clipboard
| Challenge: | Existing efforts to conventionalize the dialectal orthography of Arabic have focused on specific dialects and made ad hoc decisions. |
| Approach: | They propose a set of guidelines and meta-guidelines for conventional orthography of Arabic dialects . they apply them to 28 Arab city dialects from Rabat to Muscat . |
| Outcome: | The proposed guidelines and resources are being used by three large Arabic dialect processing projects in three universities. |
Copied to clipboard
| Challenge: | Existing dependency parsing models for Arabic use complementary annotations, CATiB and UD treebanks, and partially created trees for one annotation are also available to the other as features for the score function. |
| Approach: | They propose to use Arabic dependency annotations to parse projective dependency trees using CATiB and UD treebanks. |
| Outcome: | The proposed model gives 9.9% error reduction on CATiB and 6.1% on UD compared to a strong baseline and ablation tests show that the main contribution is given by sharing tree representation between tasks, and not simply sharing biLSTM layers as is often performed in NLP multitask systems. |
Copied to clipboard
| Challenge: | Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA). |
| Approach: | They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use. |
| Outcome: | The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety. |
Copied to clipboard
| Challenge: | evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets. |
| Approach: | They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries . |
| Outcome: | The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions . |
Copied to clipboard
| Challenge: | a large-scale leveled readability lexicon for Modern Standard Arabic is not available in many other languages. |
| Approach: | They propose a large-scale leveled readability lexicon for Arabic with 26,000 lemmas . they manually annotate a lexico from three different regions in the arab world . |
| Outcome: | The proposed lexicon is publicly available for Arabic readability tasks. |
Copied to clipboard
| Challenge: | Using a reading corpus in Modern Standard Arabic, we explore the lexical coverage of textbooks and unabridged works of fiction. |
| Approach: | They propose to use textbooks from the United Arab Emirates curriculum and a reading corpus in Modern Standard Arabic to enrich the sparse collection of resources available for educational applications. |
| Outcome: | The corpus spans all 12 grades and contains 129 unabridged works of fiction spanning grades 1-12 . lexical coverage is compared to other genres, and the results show that the two sub-corpora are similar to each other to measure their genres. |
Copied to clipboard
| Challenge: | Standard Arabic morphology is rich, but Arabic dialects introduce more complexity. |
| Approach: | They propose a joint morphological annotation and spelling correction system for Arabic texts . they propose morphology tools that can be used to help with productivity . |
| Outcome: | The proposed system is based on a standard and dialectal Arabic text. |
Copied to clipboard
| Challenge: | Previous work on Arabic Dialect identification focused on specific dialect levels and labels . since dialectal differences tend to be more subtle relative terms to language differences, the DID task is harder than language identification. |
| Approach: | They propose to define a standard hierarchical schema for Arabic Dialect identification . they map 29 different data sets to this schema and use it to aggregate the data . |
| Outcome: | The proposed schemas and methods are extensible to other languages and dialect groups. |
Copied to clipboard
| Challenge: | Despite advances in the field of natural language processing, many dialectal Arabic varieties are lagging behind . despite advances in NLP, many Arabic dialects are considered under-resourced . |
| Approach: | They propose a full morphological analysis and disambiguation system for Gulf Arabic . they use existing state-of-the-art morphology tools to investigate the effects of different data sizes and combinations of morphologists. |
| Outcome: | The proposed system improves on the existing system for Arabic . it is based on a set of data sizes and combinations of morphological analyzers . |
Copied to clipboard
| Challenge: | Palmyra 3.0 is a cloud-based platform for morphology and syntax annotation. |
| Approach: | They present Palmyra 3.0, a cloud-based platform for morphology and syntax annotation. |
| Outcome: | Palmyra 3.0 provides configuration files for a number of predefined formalisms, such as UD and CATiB, and a variety of user-friendly features to support annotators. |
Copied to clipboard
| Challenge: | Existing work on Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic at most (6-way classification). |
| Approach: | They propose to tackle a fine-grained Arabic dialect classification task covering 25 cities from across the Arab World, in addition to Standard Arabic. |
| Outcome: | The proposed task can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words and reach more than 90% when we consider 16 words. |
Copied to clipboard
| Challenge: | Recent work typically frames morphophonology as generating surface forms from abstract underlying representations (URs) this theory-laden assumption is expensive to annotate, especially in low-resource settings. |
| Approach: | a new approach frames morphophonology as generating surface forms from abstract underlying representations by applying phonological rules or constraints. |
| Outcome: | The proposed model removes the need to posit or label URs and lets the model exploit the surface evidence directly. |
Copied to clipboard
| Challenge: | Existing efforts to assess the readability of Arabic text are limited due to its rich morphology, complex syntax, and ambiguous orthography. |
| Approach: | They propose a web-based system for fine-grained, sentence-level Arabic readability assessment. |
| Outcome: | The demo provides two main functionalities for educators, content creators, language learners, and researchers: (1) a Search interface to explore the annotated dataset for text selection and resource development; (2) an Analyze interface to assign detailed readability labels to Arabic texts at the sentence level. |
Copied to clipboard
| Challenge: | Code-switching (CSW) text generation is a popular solution to address data scarcity. |
| Approach: | They compare linguistic theories, lexical replacements and back-translation approaches to Egyptian Arabic-English CSW. |
| Outcome: | The proposed methods perform best on machine translation and quality evaluation. |
Copied to clipboard
| Challenge: | Demo paper describes a web-based system for automatic dialect identification for Arabic text. |
| Approach: | They present a web-based system for automatic dialect identification for Arabic text that distinguishes between 25 Arab cities and Modern Standard Arabic. |
| Outcome: | The proposed system distinguishes among the dialects of 25 Arab cities (from Rabat to Muscat) and Modern Standard Arabic (MSA). |
Copied to clipboard
| Challenge: | CAMeL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis. |
| Approach: | They present CAMeL Tools, an open-source Python toolkit for Arabic natural language processing . CAMeleL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis. |
| Outcome: | The proposed tools are based on CAMeL Tools, an open-source Python toolkit for Arabic natural language processing. |
Copied to clipboard
| Challenge: | Our corpus includes 159K words selected from 15 publicly available Arabic fiction novels . text simplification aims to reduce the complexity of a text while maintaining the overall grammaticality and core content. |
| Approach: | They propose to annotate Arabic parallel corpus for text simplification targeting school-aged learners. |
| Outcome: | The SAMER Corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels. |
Copied to clipboard
| Challenge: | a paradigm discovery problem is a task of learning an inflectional morphological system from unannotated sentences. |
| Approach: | They formalize the paradigm discovery problem and develop evaluation metrics for judging systems . they use word embeddings and string similarity to cluster forms by cell and by paradigm . |
| Outcome: | The proposed system suggests clustering by cell across different inflection classes is the most pressing challenge for future work. |
Copied to clipboard
| Challenge: | Recent years have brought about very fast developments in Natural Language Processing (NLP), but many other languages are overlooked due to limited resources. |
| Approach: | They propose to repurpose a multilingual BELEBELE dataset for a task of extractive QA in the style of machine reading comprehension. |
| Outcome: | The proposed approach could be used to extract QA in the style of machine reading comprehension. |
Copied to clipboard
| Challenge: | The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages. |
| Approach: | They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. |
| Outcome: | The proposed schema has added 66 new languages, including 24 endangered languages. |
Copied to clipboard
| Challenge: | Thousands of JA texts are available online, covering genres such as philosophy, biblical commentary, and Bible translations. |
| Approach: | They propose a two-step approach to automatically transliterate Judeo-Arabic into Arabic script using simple character-level mapping followed by post-correction to address grammatical and orthographic errors. |
| Outcome: | The proposed method enables Arabic NLP tools to perform morphosyntactic tagging and machine translation, which would have been impossible on the original texts. |
Copied to clipboard
| Challenge: | Text editing is a wellstudied problem for grammatical error correction (GEC) but it is not the most efficient for morphologically rich languages like Arabic. |
| Approach: | They propose a text editing approach that derives edit tags directly from data, eliminating the need for language-specific edits. |
| Outcome: | The proposed approach achieves SOTA results on Arabic and performs on par with SOTA on two other languages. |
Copied to clipboard
| Challenge: | Recent advances in text normalization have limited applications in other languages . a novel approach to text normalizing uses character embeddings and word embedds . |
| Approach: | They propose a sequence-to-sequence model with character-based attention that uses pre-trained word embeddings to model subword information. |
| Outcome: | The proposed model achieves state-of-the-art F1 score on Arabic spelling correction task despite being small and unsuited for the task. |
Copied to clipboard
| Challenge: | a novel character-level sequence-to-sequence lemmatization model uses generic n-gram embeddings to map word/lemma pairs . semitic languages, like Arabic and Hebrew, add other challenges to handle unseen words . |
| Approach: | They propose a character-level sequence-to-sequence lemmatization model . they use generic n-gram embeddings, concatenative (stems) and templatic (roots and patterns) morphological subwords. |
| Outcome: | The proposed model outperforms other linguistically-driven models with generic n-gram embeddings . the best system handles word/lemma pairs that are both unseen in the training data . |
Copied to clipboard
| Challenge: | Existing tagsets for Arabic are difficult to combine due to the diversity of their features. |
| Approach: | They propose an Arabic Extended Morphological Analysis and Disambiguation Tagset which facilitates conversion and unification of Arabic tagsets. |
| Outcome: | The proposed tagset facilitates conversion and unification of different tagsetes used to annotate Arabic corpora. |
Copied to clipboard
| Challenge: | a small minority of dictionaries specify the readability level of their words, let alone their lexical relations with other words. |
| Approach: | They propose to use Arabic lemmas, roots, English glosses, related Arabic words and phrases to provide a readability leveled Arabic thesaurus interface. |
| Outcome: | The proposed system provides the user with lemmas, roots, English glosses, related Arabic words and phrases, and readability on a five-level readability scale. |
Copied to clipboard
| Challenge: | Modern Standard Arabic (MSA) nominals present many morphological and lexical modeling challenges that have not been consistently addressed before. |
| Approach: | They propose to use a morphological framework to model Arabic nominals using a proposed morphology framework. |
| Outcome: | The proposed model improves accuracy and consistency compared to a commonly used morphological analyzer and generator. |
Copied to clipboard
| Challenge: | PALMYRA is an annotation tool designed to help with syntactic annotation of morphologically rich languages. |
| Approach: | They present PALMYRA, a platform independent graphical dependency tree visualization and editing software. |
| Outcome: | PALMYRA is an graphical dependency tree visualization and editing software designed to support syntactic annotation of morphologically rich languages. |
Copied to clipboard
| Challenge: | Time-Offset Interaction Applications (TOIAs) simulate face-to-face conversations between humans and digital human avatars recorded in the past. |
| Approach: | They propose a methodology for creating the knowledge base for a TOIA, a dialogue corpus, and baselines for single-turn answer retrieval. |
| Outcome: | The proposed method lets the avatar maker list pairs by intuition, guessing what possible questions a user may ask to the avatar. |
Copied to clipboard
| Challenge: | Existing studies on grammatical error correction (GEC) in morphologically rich languages have been limited due to data scarcity and language complexity. |
| Approach: | They propose to use Arabic GEC to improve performance across three datasets . they define Arabic grammatical error detection task as auxiliary input . |
| Outcome: | The proposed models achieve SOTA results on two Arabic GEC shared task datasets and establish a strong benchmark on a recently created dataset. |
Copied to clipboard
| Challenge: | Arabic dialects are non-standard varieties of Arabic commonly spoken across the Arab world, but lack standard orthographies. |
| Approach: | They present a corpus of 10,000 sentences from five Arabic city dialects represented in the Conventional Orthography for Dialectal Arabic (CODA) they use a bootstrapping technique to speed up annotation and compare similarity between dialects before and after CODA annotation. |
| Outcome: | The proposed method speeds up the annotation process and shows similarity between the dialects before and after CODA annotation. |
Copied to clipboard
| Challenge: | Morphological tagging is challenging for morphologically rich languages due to the large combined target space and the need for more training data to minimize model sparsity. |
| Approach: | They propose to use multitask learning and adversarial training to address morphological richness and dialectal variations in the context of full morphology. |
| Outcome: | The proposed model achieves state-of-the-art for two dialectal variants: Modern Standard Arabic (high-resource “dialect”) and Egyptian Arabic (low-resourced dialect). |