Papers by Arif Khan

5 papers
BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla Lemmatizer (2023.findings-emnlp)

Copied to clipboard

Challenge: Lemmatization holds significance in both natural language processing (NLP) and linguistics due to the highly inflected nature and morphological richness of Bangla text.
Approach: They propose linguistic rules for lemmatization and utilize a dictionary along with the rules to design a lemma specifically for Bangla.
Outcome: The proposed system achieves 96.36% accuracy when tested against a manually annotated test dataset.
BLISS: An Agent for Collecting Spoken Dialogue Data about Health and Well-being (2020.lrec-1)

Copied to clipboard

Challenge: Structured interviews are a time-consuming and inefficient way to gather information about people's well-being.
Approach: They propose to build an artificial intelligence agent which asks questions about happiness . they build a prototype of the agent and collect 55 spoken dialogues .
Outcome: The proposed agent collects 55 spoken dialogues and asks users about happiness and well-being.
WER We Stand: Benchmarking Urdu ASR Models (2025.coling-main)

Copied to clipboard

Challenge: This paper analyzes the performance of three ASR models for low-resource languages like Urdu . low-rural languages like urdu have significant gaps in accuracy and reliability .
Approach: They evaluate the performance of three ASR models: Whisper, MMS, and Seamless-M4T . they present the first conversational speech dataset for benchmarking Urdu ASR systems .
Outcome: The proposed model families outperform Whisper, MMS, and Seamless-M4T on two types of speech datasets.
Kahaani: A Multimodal Co-Creative Storytelling System (2026.eacl-srw)

Copied to clipboard

Challenge: Kahaani is a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to address the challenge of sustaining engagement to foster educational narrative experiences.
Approach: They propose a multimodal, co-creative storytelling system that leverages Generative Artificial Intelligence to help children develop their storytelling skills.
Outcome: The proposed system combines large language models, text-to-speech, and music generation to produce a rich, immersive, and accessible storytelling experience.
A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks (L18-1)

Copied to clipboard

Challenge: Phonetic segmentation is the process of splitting speech into distinct phonetic units . methods for automatic segmentation are not always accurate enough .
Approach: They propose to model phonetic segmentation as close as possible to manual segmentation by recording experts performing a segmentation task.
Outcome: This corpus captures human segmentation behavior by recording experts performing a segmentation task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations