Papers by Samee Arif
Generalists vs. Specialists: Evaluating Large Language Models for Urdu (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Urdu is underrepresented in natural language processing, yet it is underserved. |
| Approach: | They compare general-purpose models with special-purpose ones that have been fine-tuned on specific tasks. |
| Outcome: | The proposed models outperform general-purpose models on seven classification and seven generation tasks. |
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)
Copied to clipboard
| Challenge: | This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability . |
| Approach: | They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems . |
| Outcome: | The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets. |
Kahaani: A Multimodal Co-Creative Storytelling System (2026.eacl-srw)
Copied to clipboard
| Challenge: | Kahaani is a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to address the challenge of sustaining engagement to foster educational narrative experiences. |
| Approach: | They propose a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to help children develop their storytelling skills. |
| Outcome: | The proposed system combines large language models, text-to-speech, and music generation to produce a rich, immersive, and accessible storytelling experience. |
UQA: Corpus for Urdu Question Answering (2024.lrec-main)
Copied to clipboard
| Challenge: | Urdu is a low-resource language with over 70 million native speakers . expanding the reach of NLP to languages other than English is crucial for advancing multilingual AI systems. |
| Approach: | They introduce a novel dataset for question answering and text comprehension in Urdu . they use a technique called EATS which preserves the answer spans in translated context paragraphs . |
| Outcome: | The proposed dataset preserves answer spans in translated context paragraphs. |