Papers by Arif Khan
BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla Lemmatizer (2023.findings-emnlp)
Copied to clipboard
Sadia Afrin, Md. Shahad Mahmud Chowdhury, Md. Islam, Faisal Khan, Labib Chowdhury, Md. Mahtab, Nazifa Chowdhury, Massud Forkan, Neelima Kundu, Hakim Arif, Mohammad Mamun Or Rashid, Mohammad Amin, Nabeel Mohammed
| Challenge: | Lemmatization holds significance in both natural language processing (NLP) and linguistics due to the highly inflected nature and morphological richness of Bangla text. |
| Approach: | They propose linguistic rules for lemmatization and utilize a dictionary along with the rules to design a lemma specifically for Bangla. |
| Outcome: | The proposed system achieves 96.36% accuracy when tested against a manually annotated test dataset. |
BLISS: An Agent for Collecting Spoken Dialogue Data about Health and Well-being (2020.lrec-1)
Copied to clipboard
Jelte van Waterschoot, Iris Hendrickx, Arif Khan, Esther Klabbers, Marcel de Korte, Helmer Strik, Catia Cucchiarini, Mariët Theune
| Challenge: | Structured interviews are a time-consuming and inefficient way to gather information about people's well-being. |
| Approach: | They propose to build an artificial intelligence agent which asks questions about happiness . they build a prototype of the agent and collect 55 spoken dialogues . |
| Outcome: | The proposed agent collects 55 spoken dialogues and asks users about happiness and well-being. |
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)
Copied to clipboard
| Challenge: | This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability . |
| Approach: | They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems . |
| Outcome: | The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets. |
Kahaani: A Multimodal Co-Creative Storytelling System (2026.eacl-srw)
Copied to clipboard
| Challenge: | Kahaani is a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to address the challenge of sustaining engagement to foster educational narrative experiences. |
| Approach: | They propose a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to help children develop their storytelling skills. |
| Outcome: | The proposed system combines large language models, text-to-speech, and music generation to produce a rich, immersive, and accessible storytelling experience. |
A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks (L18-1)
Copied to clipboard
| Challenge: | Phonetic segmentation is the process of splitting speech into distinct phonetic units . methods for automatic segmentation are not always accurate enough . |
| Approach: | They propose to model phonetic segmentation as close as possible to manual segmentation by recording experts performing a segmentation task. |
| Outcome: | This corpus captures human segmentation behavior by recording experts performing a segmentation task. |