Papers by Samee Arif

4 papers
Generalists vs. Specialists: Evaluating Large Language Models for Urdu (2024.findings-emnlp)

Copied to clipboard

Challenge: Urdu is underrepresented in natural language processing, yet it is underserved.
Approach: They compare general-purpose models with special-purpose ones that have been fine-tuned on specific tasks.
Outcome: The proposed models outperform general-purpose models on seven classification and seven generation tasks.
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)

Copied to clipboard

Challenge: This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability .
Approach: They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems .
Outcome: The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets.
Kahaani: A Multimodal Co-Creative Storytelling System (2026.eacl-srw)

Copied to clipboard

Challenge: Kahaani is a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to address the challenge of sustaining engagement to foster educational narrative experiences.
Approach: They propose a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to help children develop their storytelling skills.
Outcome: The proposed system combines large language models, text-to-speech, and music generation to produce a rich, immersive, and accessible storytelling experience.
UQA: Corpus for Urdu Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: Urdu is a low-resource language with over 70 million native speakers . expanding the reach of NLP to languages other than English is crucial for advancing multilingual AI systems.
Approach: They introduce a novel dataset for question answering and text comprehension in Urdu . they use a technique called EATS which preserves the answer spans in translated context paragraphs .
Outcome: The proposed dataset preserves answer spans in translated context paragraphs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations