Papers by Abdelrahman Sadallah
Commonsense Reasoning in Arab Culture (2025.acl-long)
Copied to clipboard
Abdelrahman Sadallah, Junior Cedric Tonga, Khalid Almubarak, Saeed Almheiri, Farah Atif, Chatrine Qwaider, Karima Kadaoui, Sara Shatnawi, Yaser Alesh, Fajri Koto
| Challenge: | Existing studies on commonsense reasoning in Arabic have relied on machine translations that lack cultural depth and introduce anglocentric biases. |
| Approach: | They propose a commonsense reasoning dataset in Arabic that covers 13 Arab countries. |
| Outcome: | The proposed dataset covers 13 countries across the Gulf, Levant, North Africa, and the Nile Valley. |
Instruction-Guided Poetry Generation in Arabic and Its Dialects (2026.findings-acl)
Copied to clipboard
Abdelrahman Sadallah, Kareem Elozeiri, Mervat Abassy, Rania Elbadry, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Fajri Koto
| Challenge: | Existing literature on Arabic poetry has focused on analysis tasks such as interpretation or metadata prediction, e.g., rhyme schemes and titles. |
| Approach: | They propose to use a large-scale instruction-based dataset to generate Arabic poetry based on predefined criteria such as style and rhyme . |
| Outcome: | The proposed model can generate poetry that is aligned with user requirements, based on automated metrics and human evaluation with native Arabic speakers. |
The Good, the Bad and the Constructive: Automatically Measuring Peer Review’s Utility for Authors (2025.emnlp-main)
Copied to clipboard
| Challenge: | Providing constructive feedback to authors is a core component of peer review . authors lack guidance on how to improve their review, a problem that is often overlooked . |
| Approach: | They use a RevUtil dataset to benchmark fine-tuned models for assessing review comments . they find that machine-generated reviews generally underperform human reviews on these aspects . |
| Outcome: | The proposed model outperforms closed models on four aspects of review comments . the proposed model achieves agreement levels comparable to and exceeding those of human models . |
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic (2024.findings-acl)
Copied to clipboard
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, Timothy Baldwin
| Challenge: | evaluating language models in Arabic remains challenging due to limited datasets . focus has shift to reasoning and knowledge-intensive tasks due to lack of relevant datasets. |
| Approach: | They propose to use ArabicMMLU to evaluate models' understanding of Arabic . they use 40 tasks and 14,575 multiple-choice questions from school exams in different countries . |
| Outcome: | The ArabicMMLU is the first multi-task language understanding benchmark for the Arabic language . it is based on 40 tasks and 14,575 multiple-choice questions in modern standard Arabic . the models are based in different countries across North Africa, the Levant, and the Gulf regions . |
What Makes Cryptic Crosswords Challenging for LLMs? (2025.coling-main)
Copied to clipboard
| Challenge: | Recent research suggests that solving cryptic crosswords is challenging even for modern NLP models, including Large Language Models (LLMs). |
| Approach: | They establish benchmark results for three popular LLMs: Gemma2, LLaMA3 and ChatGPT, and investigate why these models struggle to achieve superior performance. |
| Outcome: | The proposed models perform significantly below humans on the cryptic crossword puzzle task, while human solvers achieve 99% accuracy. |