Papers by Bashar Alhafni

11 papers
Do Diacritics Matter? Evaluating the Impact of Arabic Diacritics on Tokenization and LLM Benchmarks (2026.findings-eacl)

Copied to clipboard

Challenge: Diacritics can significantly influence language processing tasks in Arabic . their presence can increase subword fragmentation during tokenization, reducing performance .
Approach: They analyze the impact of diacritics on tokenization and benchmark task performance across major Large Language Models.
Outcome: The proposed model is robust to diacritics, but full diacritization leads to token fragmentation and degraded performance.
mEdIT: Multilingual Text Editing via Instruction Tuning (2024.naacl-long)

Copied to clipboard

Challenge: mEdIT is a multi-lingual extension to CoEdit for writing assistance.
Approach: They propose to train multi-lingual large language models (LLMs) by fine-tuning them via instruction tuning.
Outcome: The proposed model performs well on multilingual text editing benchmarks and generalizes well to new languages.
The Arabic Parallel Gender Corpus 2.0: Extensions and Analyses (2022.lrec-1)

Copied to clipboard

Challenge: Gender bias in natural language processing (NLP) applications has been receiving increasing attention, largely due to the lack of datasets and resources.
Approach: They propose a corpus for gender identification and rewriting in contexts involving one or two target users with independent grammatical gender preferences.
Outcome: The proposed corpus expands on Habash et al.'s Arabic Parallel Gender Corpus (APGC) by adding second person targets and increasing the total number of sentences over 6.5 times, reaching over 590K words.
User-Centric Gender Rewriting (2022.naacl-main)

Copied to clipboard

Challenge: Existing systems that embed and amplify gender bias can still exhibit and exacerbate this problem.
Approach: They propose a multi-step system that combines the positive aspects of rule-based and neural rewriting models to provide personalized outputs based on the users’ grammatical gender preferences.
Outcome: The proposed system achieves 88.42 M2 F0.5 on a blind test set and improves over previous work on the first-person-only version of this task by 3.05 absolute increase in M2F0.5.
Arabic Word-level Readability Visualization for Assisted Text Simplification (2022.emnlp-demos)

Copied to clipboard

Challenge: a Google Docs add-on for automatic Arabic word-level readability visualization is available for free.
Approach: They propose a Google Docs add-on for automatic Arabic word-level readability visualization.
Outcome: The proposed add-on can be used to assess the reading difficulty of a text and identify difficult words as part of manual text simplification.
CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing (2020.lrec-1)

Copied to clipboard

Challenge: CAMeL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis.
Approach: They present CAMeL Tools, an open-source Python toolkit for Arabic natural language processing . CAMeleL Tools provides utilities for pre-processing, morphological modeling, Dialect Identification, Named Entity Recognition and sentiment analysis.
Outcome: The proposed tools are based on CAMeL Tools, an open-source Python toolkit for Arabic natural language processing.
The SAMER Arabic Text Simplification Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Our corpus includes 159K words selected from 15 publicly available Arabic fiction novels . text simplification aims to reduce the complexity of a text while maintaining the overall grammaticality and core content.
Approach: They propose to annotate Arabic parallel corpus for text simplification targeting school-aged learners.
Outcome: The SAMER Corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels.
CrisisLTLSum: A Benchmark for Local Crisis Event Timeline Extraction and Summarization (2022.findings-emnlp)

Copied to clipboard

Challenge: Timeline extraction and abstractive summarization are critical tasks for leveraging large numbers of social media posts about events.
Approach: They propose to build a semi-automated cluster-then-refine algorithm to extract local crisis event timelines from Twitter.
Outcome: The proposed approach performs better than human models on extraction and summarization tasks.
A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic (2026.eacl-long)

Copied to clipboard

Challenge: Thousands of JA texts are available online, covering genres such as philosophy, biblical commentary, and Bible translations.
Approach: They propose a two-step approach to automatically transliterate Judeo-Arabic into Arabic script using simple character-level mapping followed by post-correction to address grammatical and orthographic errors.
Outcome: The proposed method enables Arabic NLP tools to perform morphosyntactic tagging and machine translation, which would have been impossible on the original texts.
Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study (2025.acl-long)

Copied to clipboard

Challenge: Text editing is a wellstudied problem for grammatical error correction (GEC) but it is not the most efficient for morphologically rich languages like Arabic.
Approach: They propose a text editing approach that derives edit tags directly from data, eliminating the need for language-specific edits.
Outcome: The proposed approach achieves SOTA results on Arabic and performs on par with SOTA on two other languages.
Advancements in Arabic Grammatical Error Detection and Correction: An Empirical Investigation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on grammatical error correction (GEC) in morphologically rich languages have been limited due to data scarcity and language complexity.
Approach: They propose to use Arabic GEC to improve performance across three datasets . they define Arabic grammatical error detection task as auxiliary input .
Outcome: The proposed models achieve SOTA results on two Arabic GEC shared task datasets and establish a strong benchmark on a recently created dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations