Papers by Nizar Habash

64 papers
The Effectiveness of Simple Hybrid Systems for Hypernym Discovery (P19-1)

Copied to clipboard

Challenge: Recent work utilizing a mix of pattern-based and distributional approaches have yielded state-of-the-art results on two domain-specific English hypernym discovery tasks.
Approach: They evaluate the contribution of pattern-based and distributional approaches to hybrid modeling by evaluating baseline models from each paradigm.
Outcome: The proposed approach outperforms all non-hybrid approaches on two domain-specific English hypernym discovery tasks and outperformed other approaches.
Do Diacritics Matter? Evaluating the Impact of Arabic Diacritics on Tokenization and LLM Benchmarks (2026.findings-eacl)

Copied to clipboard

Challenge: Diacritics can significantly influence language processing tasks in Arabic . their presence can increase subword fragmentation during tokenization, reducing performance .
Approach: They analyze the impact of diacritics on tokenization and benchmark task performance across major Large Language Models.
Outcome: The proposed model is robust to diacritics, but full diacritization leads to token fragmentation and degraded performance.
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment (2025.findings-acl)

Copied to clipboard

Challenge: Texts above a student's readability level can lead to disengagement and disengagement . Developing readability models is crucial for improving literacy, language learning, and academic performance.
Approach: They introduce the Balanced Arabic Readability Evaluation Corpus (BAREC) a large-scale, fine-grained dataset for Arabic readability assessment.
Outcome: The proposed model outperforms existing methods in Arabic readability assessment.
Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data (2025.findings-emnlp)

Copied to clipboard

Challenge: Maltese is a Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English.
Approach: They investigate whether Arabic-language resources can support Maltese natural language processing . they introduce transliteration schemes and machine translation approaches to align Arabic text with Maltesen .
Outcome: The proposed techniques can significantly improve Maltese natural language processing tasks.
Arabic Natural Language Processing (2022.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial provides background information for system developers and researchers working with Arabic in its various forms.
Approach: This tutorial provides the necessary background information for working with Arabic in its various forms.
Outcome: This tutorial will explain various Arabic linguistic phenomena and review the state-of-the-art in Arabic processing.
The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work addresses this issue by modeling dialectness as a continuous variable . however, ALDi reduces complex variation to a single dimension .
Approach: They propose a way to model Arabic dialectness as a continuous variable . they propose etymology-aware edit distance and a regression model to model AGS .
Outcome: The proposed approach outperforms baselines on a multi-dialect benchmark.
Exploring Segmentation Approaches for Neural Machine Translation of Code-Switched Egyptian Arabic-English Text (2023.eacl-main)

Copied to clipboard

Challenge: Code-switching (CS) is a problem in machine translation, but its performance is not investigated for CS settings.
Approach: They propose to use morphological segmentation techniques for machine translation tasks . they compare morphology-based and frequency-based segmentation for MT tasks based on data size .
Outcome: The proposed approach performs best in MT tasks but under-performs in other languages.
Joint Diacritization, Lemmatization, Normalization, and Fine-Grained Morphological Tagging (2020.acl-main)

Copied to clipboard

Challenge: a word can have multiple interpretations and is one of many inflected forms of the same concept or lemma.
Approach: They propose to model morphological features jointly, whether lexicalized or non-lexicalised . their results are compared to Arabic and Egyptian Arabic .
Outcome: The proposed model achieves 20% relative error reduction in Arabic and 11% in Egyptian Arabic.
ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus (2022.lrec-1)

Copied to clipboard

Challenge: ZAEBUC is an annotated Arabic-English bilingual writer corpus . it is a corpus of short essays written by first-year university students .
Approach: They propose to use a standard Arabic-English bilingual writer corpus to match comparable texts written by the same writer on different occasions.
Outcome: The ZAEBUC corpus is an annotated Arabic-English bilingual writer corpus by first-year university students at Zayed University in the United Arab Emirates.
M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection (2024.eacl-long)

Copied to clipboard

Challenge: Large language models generate fluent responses to user queries, but they are also susceptible to misuse in journalism, education, and academia.
Approach: They propose a large-scale benchmark for machine-generated text detection that is a multi-generator, multi-domain, and multi-lingual corpus.
Outcome: The proposed system can detect machine-generated text and pinpoint misuse . the proposed system is based on a large-scale benchmark dataset .
A Cross-lingual Messenger with Keyword Searchable Phrases for the Travel Domain (C18-2)

Copied to clipboard

Challenge: Query Translator is a cross-lingual messaging app for the travel domain that automatically translates conversations . the application addresses common cross-linguistic communication issues such as translation accuracy, speed, privacy and personalization.
Approach: They present a cross-lingual messaging app that automatically translates conversations while supporting keyword-to-sentence matching.
Outcome: The proposed app translates conversations while supporting keyword-to-sentence matching.
CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing (L18-1)

Copied to clipboard

Challenge: Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks.
Approach: They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages.
Outcome: The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks.
Morphosyntactic Tagging with Pre-trained Language Models for Arabic and its Dialects (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained morphosyntactic tagging models outperform existing systems in Modern Standard Arabic and all the Arabic dialects studied.
Approach: They present results on morphosyntactic tagging across different varieties of Arabic using pre-trained transformer language models.
Outcome: The proposed models outperform existing systems in Modern Standard Arabic, 2.8% in Gulf, 1.6% in Egyptian, and 8.3% in Levantine.
A Parallel Corpus of Arabic-Japanese News Articles (L18-1)

Copied to clipboard

Challenge: a large-scale parallel corpora with manually verified subsets of sentences has been used for machine translation between major language pairs.
Approach: They describe the creation process and statistics of the Arabic-Japanese portion of the TUFS Media Corpus . they also report the first results of Arabic-japanese phrase-based machine translation trained on the corpus based on the Arabic corpus.
Outcome: The proposed corpus is a document-level parallel corpus and sentence-level parser corpus . it is the first time that Arabic-Japanese translations have been trained on it .
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)

Copied to clipboard

Challenge: CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses)
Approach: They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic.
Outcome: The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses)
Addressing Noise in Multidialectal Word Embeddings (P18-2)

Copied to clipboard

Challenge: Dialectal Arabic (DA) is problematically noisy and lacks a large corpus of non-noisy words.
Approach: They propose to use word embedding tools to maximize the informative content leveraged in each training sentence and analyze methods for representing disparate dialects in one embeddable space.
Outcome: The proposed methods improve performance on low and high frequency words while preserving accuracy on low frequency forms.
M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have brought an unprecedented surge in machine-generated text (MGT) societal implications are posed by their potential misuse and lack of training data.
Approach: They propose a benchmark to detect machine-generated text in multiple languages . they use multi-domain and multi-generator corpus to identify which model generated the text .
Outcome: The proposed benchmark compares a multilingual, multi-domain and multi-generator corpus of MGTs with human-generated content.
The Arabic Parallel Gender Corpus 2.0: Extensions and Analyses (2022.lrec-1)

Copied to clipboard

Challenge: Gender bias in natural language processing (NLP) applications has been receiving increasing attention, largely due to the lack of datasets and resources.
Approach: They propose a corpus for gender identification and rewriting in contexts involving one or two target users with independent grammatical gender preferences.
Outcome: The proposed corpus expands on Habash et al.'s Arabic Parallel Gender Corpus (APGC) by adding second person targets and increasing the total number of sentences over 6.5 times, reaching over 590K words.
Camel Morph MSA: A Large-Scale Open-Source Morphological Analyzer for Modern Standard Arabic (2024.lrec-main)

Copied to clipboard

Challenge: Camel Morph MSA is the largest open-source Modern Standard Arabic morphological analyzer and generator.
Approach: They present Camel Morph MSA, the largest open-source Arabic morphological analyzer and generator.
Outcome: The analysis can produce 1.45B analyses and 535M unique diacritizations, almost an order of magnitude larger than SAMA on a 10B word corpus.
The MADAR Arabic Dialect Corpus and Lexicon (L18-1)

Copied to clipboard

Challenge: Using a corpus of 25 Arabic city dialects and a lexicon of 1,045 concepts, we study 25 cities in a travel domain . focus on cities opens new avenues for research from dialectology to dialect identification and machine translation.
Approach: They present two Arabic language resources that are part of the Multi Arabic Dialect Applications and Resources project.
Outcome: The proposed resources are the first of their kind in terms of their coverage and fine granularity.
The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)

Copied to clipboard

Challenge: Various corpora of various sizes and representing different genres, have been created for various Arabic dialects.
Approach: They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files.
Outcome: The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP .
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of multilingual Arabic-English speech is presented in a new paper . a major bottleneck is the lack of data needed for training NLP models .
Approach: They propose a multilingual multidialectal Arabic-English speech corpus with a set of guidelines for automatic speech recognition.
Outcome: The proposed corpus includes two languages with Arabic and English spoken in multiple variants and Arabic and Arabic with various accents.
User-Centric Gender Rewriting (2022.naacl-main)

Copied to clipboard

Challenge: Existing systems that embed and amplify gender bias can still exhibit and exacerbate this problem.
Approach: They propose a multi-step system that combines the positive aspects of rule-based and neural rewriting models to provide personalized outputs based on the users’ grammatical gender preferences.
Outcome: The proposed system achieves 88.42 M2 F0.5 on a blind test set and improves over previous work on the first-person-only version of this task by 3.05 absolute increase in M2F0.5.
Cross-Lingual Transfer from Related Languages: Treating Low-Resource Maltese as Multilingual Code-Switching (2024.eacl-long)

Copied to clipboard

Challenge: Multilingual models exhibit impressive cross-lingual transfer capabilities on unseen languages, but performance is impacted when there is a script disparity with the languages used in the model’s pre-training data.
Approach: They propose a novel method to align a resource-rich language's script with a target language and train a classifier that can make informed decisions regarding the appropriate processing of each token.
Outcome: The proposed model can be used to transfer a language's scripts across multiple languages, but it is suboptimal for mixed languages, where only a subset benefits while the rest is impeded.
Lemmatization as a Classification Task: Results from Arabic across Multiple Genres (2025.emnlp-main)

Copied to clipboard

Challenge: Existing tools for lemmatization in morphologically rich languages with ambiguous orthography face inconsistent standards and limited genre coverage.
Approach: They propose two new approaches that frame lemmatization as classification into a Lemma-POS-Gloss tagset, leveraging machine translation and semantic clustering.
Outcome: The proposed models perform better than existing models and are more interpretable, the authors show.
Arabic Word-level Readability Visualization for Assisted Text Simplification (2022.emnlp-demos)

Copied to clipboard

Challenge: a Google Docs add-on for automatic Arabic word-level readability visualization is available for free.
Approach: They propose a Google Docs add-on for automatic Arabic word-level readability visualization.
Outcome: The proposed add-on can be used to assess the reading difficulty of a text and identify difficult words as part of manual text simplification.
Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI (2026.acl-long)

Copied to clipboard

Challenge: Prior studies have shown that distinguishing text generated by Large Language Models from human-written text is challenging for humans and often no better than random guessing.
Approach: They conduct extensive case study to determine the upper bound of human detection accuracy.
Outcome: The findings challenge previous conclusions on human detection accuracy across languages and domains.
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)

Copied to clipboard

Challenge: Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language .
Approach: They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research .
Outcome: This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research .
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection (2024.emnlp-demo)

Copied to clipboard

Challenge: a large number of machine-generated texts are often hard to distinguish between human-written and machine-generated text . this raises concerns about potential misuse, especially within educational and academic domains .
Approach: They propose a system that can detect whether a text is human-written or machine-generated . they use a fine-grained classification schema to identify the use of machine-generated text .
Outcome: The proposed system can distinguish between human-written and machine-generated text . it can detect attempts to obfuscate the fact that a text was machine- generated .
Camelira: An Arabic Multi-Dialect Morphological Disambiguator (2022.emnlp-demos)

Copied to clipboard

Challenge: Camelira is a web-based Arabic multi-dialect morphological disambiguation tool that covers modern standard Arabic, Egyptian, Gulf, and Levantine.
Approach: They propose a web-based Arabic multi-dialect morphological disambiguation tool that covers modern standard Arabic, Egyptian, Gulf, and Levantine.
Outcome: The proposed tool covers modern standard Arabic, Egyptian, Gulf, and Levantine . it also provides an option to automatically choose an appropriate disambiguator based on the prediction of a dialect identification component.
Computational Benchmarks for Egyptian Arabic Child Directed Speech (2026.eacl-long)

Copied to clipboard

Challenge: Existing CDS corpora in English are limited due to their limited size or lack of linguistic annotation.
Approach: They propose to map Egyptian Arabic IPA tokens to Arabic script and add core part-of-speech tags and lemmas aligned with existing Arabic morphological resources.
Outcome: The proposed version of the Egyptian Arabic CHILDES corpus is available online.
Noise-Robust Morphological Disambiguation for Dialectal Arabic (N18-1)

Copied to clipboard

Challenge: Noisy content is non-canonical in nature, with lexical, orthographic, and phonetic variations.
Approach: They propose a neural morphological tagging and disambiguation model for Egyptian Arabic with various extensions to handle noisy content.
Outcome: The proposed model achieves about 5% relative error reduction over a state-of-the-art baseline for Egyptian Arabic.
Unified Guidelines and Resources for Arabic Dialect Orthography (L18-1)

Copied to clipboard

Challenge: Existing efforts to conventionalize the dialectal orthography of Arabic have focused on specific dialects and made ad hoc decisions.
Approach: They propose a set of guidelines and meta-guidelines for conventional orthography of Arabic dialects . they apply them to 28 Arab city dialects from Rabat to Muscat .
Outcome: The proposed guidelines and resources are being used by three large Arabic dialect processing projects in three universities.
Multitask Easy-First Dependency Parsing: Exploiting Complementarities of Different Dependency Representations (2020.coling-main)

Copied to clipboard

Challenge: Existing dependency parsing models for Arabic use complementary annotations, CATiB and UD treebanks, and partially created trees for one annotation are also available to the other as features for the score function.
Approach: They propose to use Arabic dependency annotations to parse projective dependency trees using CATiB and UD treebanks.
Outcome: The proposed model gives 9.9% error reduction on CATiB and 6.1% on UD compared to a strong baseline and ablation tests show that the main contribution is given by sharing tree representation between tasks, and not simply sharing biLSTM layers as is often performed in NLP multitask systems.
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)

Copied to clipboard

Challenge: Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA).
Approach: They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use.
Outcome: The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety.
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic (2024.findings-acl)

Copied to clipboard

Challenge: evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets.
Approach: They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries .
Outcome: The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions .
A Large-Scale Leveled Readability Lexicon for Standard Arabic (2020.lrec-1)

Copied to clipboard

Challenge: a large-scale leveled readability lexicon for Modern Standard Arabic is not available in many other languages.
Approach: They propose a large-scale leveled readability lexicon for Arabic with 26,000 lemmas . they manually annotate a lexico from three different regions in the arab world .
Outcome: The proposed lexicon is publicly available for Arabic readability tasks.
A Leveled Reading Corpus of Modern Standard Arabic (L18-1)

Copied to clipboard

Challenge: Using a reading corpus in Modern Standard Arabic, we explore the lexical coverage of textbooks and unabridged works of fiction.
Approach: They propose to use textbooks from the United Arab Emirates curriculum and a reading corpus in Modern Standard Arabic to enrich the sparse collection of resources available for educational applications.
Outcome: The corpus spans all 12 grades and contains 129 unabridged works of fiction spanning grades 1-12 . lexical coverage is compared to other genres, and the results show that the two sub-corpora are similar to each other to measure their genres.
MADARi: A Web Interface for Joint Arabic Morphological Annotation and Spelling Correction (L18-1)

Copied to clipboard

Challenge: Standard Arabic morphology is rich, but Arabic dialects introduce more complexity.
Approach: They propose a joint morphological annotation and spelling correction system for Arabic texts . they propose morphology tools that can be used to help with productivity .
Outcome: The proposed system is based on a standard and dialectal Arabic text.
Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification (2022.lrec-1)

Copied to clipboard

Challenge: Previous work on Arabic Dialect identification focused on specific dialect levels and labels . since dialectal differences tend to be more subtle relative terms to language differences, the DID task is harder than language identification.
Approach: They propose to define a standard hierarchical schema for Arabic Dialect identification . they map 29 different data sets to this schema and use it to aggregate the data .
Outcome: The proposed schemas and methods are extensible to other languages and dialect groups.
Morphological Analysis and Disambiguation for Gulf Arabic: The Interplay between Resources and Methods (2020.lrec-1)

Copied to clipboard

Challenge: Despite advances in the field of natural language processing, many dialectal Arabic varieties are lagging behind . despite advances in NLP, many Arabic dialects are considered under-resourced .
Approach: They propose a full morphological analysis and disambiguation system for Gulf Arabic . they use existing state-of-the-art morphology tools to investigate the effects of different data sizes and combinations of morphologists.
Outcome: The proposed system improves on the existing system for Arabic . it is based on a set of data sizes and combinations of morphological analyzers .
Palmyra 3.0: A User-Friendly Cloud-Based Platform for Morphology and Dependency Syntax Annotation (2024.lrec-main)

Copied to clipboard

Challenge: Palmyra 3.0 is a cloud-based platform for morphology and syntax annotation.
Approach: They present Palmyra 3.0, a cloud-based platform for morphology and syntax annotation.
Outcome: Palmyra 3.0 provides configuration files for a number of predefined formalisms, such as UD and CATiB, and a variety of user-friendly features to support annotators.
Fine-Grained Arabic Dialect Identification (C18-1)

Copied to clipboard

Challenge: Existing work on Arabic Dialect Identification typically targeted coarse-grained five dialect classes plus Standard Arabic at most (6-way classification).
Approach: They propose to tackle a fine-grained Arabic dialect classification task covering 25 cities from across the Arab World, in addition to Standard Arabic.
Outcome: The proposed task can identify the exact city of a speaker at an accuracy of 67.9% for sentences with an average length of 7 words and reach more than 90% when we consider 16 words.
Radical Allomorphy: Phonological Surface Forms without Phonology (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent work typically frames morphophonology as generating surface forms from abstract underlying representations (URs) this theory-laden assumption is expensive to annotate, especially in low-resource settings.
Approach: a new approach frames morphophonology as generating surface forms from abstract underlying representations by applying phonological rules or constraints.
Outcome: The proposed model removes the need to posit or label URs and lets the model exploit the surface evidence directly.
BAREC Demo: Resources and Tools for Sentence-level Arabic Readability Assessment (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing efforts to assess the readability of Arabic text are limited due to its rich morphology, complex syntax, and ambiguous orthography.
Approach: They propose a web-based system for fine-grained, sentence-level Arabic readability assessment.
Outcome: The demo provides two main functionalities for educators, content creators, language learners, and researchers: (1) a Search interface to explore the annotated dataset for text selection and resource development; (2) an Analyze interface to assign detailed readability labels to Arabic texts at the sentence level.
Data Augmentation Techniques for Machine Translation of Code-Switched Texts: A Comparative Study (2023.findings-emnlp)

Copied to clipboard

Challenge: Code-switching (CSW) text generation is a popular solution to address data scarcity.
Approach: They compare linguistic theories, lexical replacements and back-translation approaches to Egyptian Arabic-English CSW.
Outcome: The proposed methods perform best on machine translation and quality evaluation.
ADIDA: Automatic Dialect Identification for Arabic (N19-4)

Copied to clipboard

Challenge: Demo paper describes a web-based system for automatic dialect identification for Arabic text.
Approach: They present a web-based system for automatic dialect identification for Arabic text that distinguishes between 25 Arab cities and Modern Standard Arabic.
Outcome: The proposed system distinguishes among the dialects of 25 Arab cities (from Rabat to Muscat) and Modern Standard Arabic (MSA).
CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing (2020.lrec-1)

Copied to clipboard

Challenge: CAMeL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis.
Approach: They present CAMeL Tools, an open-source Python toolkit for Arabic natural language processing . CAMeleL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis.
Outcome: The proposed tools are based on CAMeL Tools, an open-source Python toolkit for Arabic natural language processing.
The SAMER Arabic Text Simplification Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Our corpus includes 159K words selected from 15 publicly available Arabic fiction novels . text simplification aims to reduce the complexity of a text while maintaining the overall grammaticality and core content.
Approach: They propose to annotate Arabic parallel corpus for text simplification targeting school-aged learners.
Outcome: The SAMER Corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels.
The Paradigm Discovery Problem (2020.acl-main)

Copied to clipboard

Challenge: a paradigm discovery problem is a task of learning an inflectional morphological system from unannotated sentences.
Approach: They formalize the paradigm discovery problem and develop evaluation metrics for judging systems . they use word embeddings and string similarity to cluster forms by cell and by paradigm .
Outcome: The proposed system suggests clustering by cell across different inflection classes is the most pressing challenge for future work.
From Multiple-Choice to Extractive QA: A Case Study for English and Arabic (2025.coling-main)

Copied to clipboard

Challenge: Recent years have brought about very fast developments in Natural Language Processing (NLP), but many other languages are overlooked due to limited resources.
Approach: They propose to repurpose a multilingual BELEBELE dataset for a task of extractive QA in the style of machine reading comprehension.
Outcome: The proposed approach could be used to extract QA in the style of machine reading comprehension.
UniMorph 4.0: Universal Morphology (2022.lrec-1)

Copied to clipboard

Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
Challenge: The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema.
Outcome: The proposed schema has added 66 new languages, including 24 endangered languages.
A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic (2026.eacl-long)

Copied to clipboard

Challenge: Thousands of JA texts are available online, covering genres such as philosophy, biblical commentary, and Bible translations.
Approach: They propose a two-step approach to automatically transliterate Judeo-Arabic into Arabic script using simple character-level mapping followed by post-correction to address grammatical and orthographic errors.
Outcome: The proposed method enables Arabic NLP tools to perform morphosyntactic tagging and machine translation, which would have been impossible on the original texts.
Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study (2025.acl-long)

Copied to clipboard

Challenge: Text editing is a wellstudied problem for grammatical error correction (GEC) but it is not the most efficient for morphologically rich languages like Arabic.
Approach: They propose a text editing approach that derives edit tags directly from data, eliminating the need for language-specific edits.
Outcome: The proposed approach achieves SOTA results on Arabic and performs on par with SOTA on two other languages.
Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models (D18-1)

Copied to clipboard

Challenge: Recent advances in text normalization have limited applications in other languages . a novel approach to text normalizing uses character embeddings and word embedds .
Approach: They propose a sequence-to-sequence model with character-based attention that uses pre-trained word embeddings to model subword information.
Outcome: The proposed model achieves state-of-the-art F1 score on Arabic spelling correction task despite being small and unsuited for the task.
Utilizing Subword Entities in Character-Level Sequence-to-Sequence Lemmatization Models (2020.coling-main)

Copied to clipboard

Challenge: a novel character-level sequence-to-sequence lemmatization model uses generic n-gram embeddings to map word/lemma pairs . semitic languages, like Arabic and Hebrew, add other challenges to handle unseen words .
Approach: They propose a character-level sequence-to-sequence lemmatization model . they use generic n-gram embeddings, concatenative (stems) and templatic (roots and patterns) morphological subwords.
Outcome: The proposed model outperforms other linguistically-driven models with generic n-gram embeddings . the best system handles word/lemma pairs that are both unseen in the training data .
EMAD: A Bridge Tagset for Unifying Arabic POS Annotations (2024.lrec-main)

Copied to clipboard

Challenge: Existing tagsets for Arabic are difficult to combine due to the diversity of their features.
Approach: They propose an Arabic Extended Morphological Analysis and Disambiguation Tagset which facilitates conversion and unification of Arabic tagsets.
Outcome: The proposed tagset facilitates conversion and unification of different tagsetes used to annotate Arabic corpora.
An Online Readability Leveled Arabic Thesaurus (2020.coling-demos)

Copied to clipboard

Challenge: a small minority of dictionaries specify the readability level of their words, let alone their lexical relations with other words.
Approach: They propose to use Arabic lemmas, roots, English glosses, related Arabic words and phrases to provide a readability leveled Arabic thesaurus interface.
Outcome: The proposed system provides the user with lemmas, roots, English glosses, related Arabic words and phrases, and readability on a five-level readability scale.
Computational Morphology and Lexicography Modeling of Modern Standard Arabic Nominals (2024.findings-eacl)

Copied to clipboard

Challenge: Modern Standard Arabic (MSA) nominals present many morphological and lexical modeling challenges that have not been consistently addressed before.
Approach: They propose to use a morphological framework to model Arabic nominals using a proposed morphology framework.
Outcome: The proposed model improves accuracy and consistency compared to a commonly used morphological analyzer and generator.
Palmyra: A Platform Independent Dependency Annotation Tool for Morphologically Rich Languages (L18-1)

Copied to clipboard

Challenge: PALMYRA is an annotation tool designed to help with syntactic annotation of morphologically rich languages.
Approach: They present PALMYRA, a platform independent graphical dependency tree visualization and editing software.
Outcome: PALMYRA is an graphical dependency tree visualization and editing software designed to support syntactic annotation of morphologically rich languages.
The Margarita Dialogue Corpus: A Data Set for Time-Offset Interactions and Unstructured Dialogue Systems (2020.lrec-1)

Copied to clipboard

Challenge: Time-Offset Interaction Applications (TOIAs) simulate face-to-face conversations between humans and digital human avatars recorded in the past.
Approach: They propose a methodology for creating the knowledge base for a TOIA, a dialogue corpus, and baselines for single-turn answer retrieval.
Outcome: The proposed method lets the avatar maker list pairs by intuition, guessing what possible questions a user may ask to the avatar.
Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on grammatical error correction (GEC) in morphologically rich languages have been limited due to data scarcity and language complexity.
Approach: They propose to use Arabic GEC to improve performance across three datasets . they define Arabic grammatical error detection task as auxiliary input .
Outcome: The proposed models achieve SOTA results on two Arabic GEC shared task datasets and establish a strong benchmark on a recently created dataset.
A Spelling Correction Corpus for Multiple Arabic Dialects (2020.lrec-1)

Copied to clipboard

Challenge: Arabic dialects are non-standard varieties of Arabic commonly spoken across the Arab world, but lack standard orthographies.
Approach: They present a corpus of 10,000 sentences from five Arabic city dialects represented in the Conventional Orthography for Dialectal Arabic (CODA) they use a bootstrapping technique to speed up annotation and compare similarity between dialects before and after CODA annotation.
Outcome: The proposed method speeds up the annotation process and shows similarity between the dialects before and after CODA annotation.
Adversarial Multitask Learning for Joint Multi-Feature and Multi-Dialect Morphological Modeling (P19-1)

Copied to clipboard

Challenge: Morphological tagging is challenging for morphologically rich languages due to the large combined target space and the need for more training data to minimize model sparsity.
Approach: They propose to use multitask learning and adversarial training to address morphological richness and dialectal variations in the context of full morphology.
Outcome: The proposed model achieves state-of-the-art for two dialectal variants: Modern Standard Arabic (high-resource “dialect”) and Egyptian Arabic (low-resourced dialect).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations