Papers by Samar Magdy

3 papers
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets.
Approach: They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding.
Outcome: The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X).
Gazelle: An Instruction Dataset for Arabic Writing Assistance (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in generative AI have transformed the landscape of writing assistance, especially through the development of Large Language Models (LLMs).
Approach: They propose to use a dataset to evaluate leading LLMs to improve their writing assistance tools in Arabic.
Outcome: The proposed dataset highlights the strengths and limitations of leading LLMs, including GPT-**4**, GPT**4o**, Cohere Command R+, and Gemini **1.5** Pro.
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: despite recent advances in speech processing, the majority of world languages and dialects remain uncovered.
Approach: They propose to collect and transcribe a new Arabic dataset for eight dialects . they also develop strong baselines exploiting the new dataset .
Outcome: The proposed dataset covers eight Arabic dialects, including Algerian, Egyptian, Emirati, Jordanian, Mauritanian, Moroccan, Palestinian, and Yemeni.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations