Papers by Shafayat Ahmed
SHONGLAP: A Large Bengali Open-Domain Dialogue Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing open-domain dialogue systems suffer from data scarcity due to unavailability of high-quality datasets for low-resource languages like Bengali. |
| Approach: | They propose to prepare large-scale open-domain dialogue datasets from podcasts and talk-shows and label them based on weak-supervision techniques. |
| Outcome: | The proposed corpus improves performance of large language models in case of downstream classification tasks during fine-tuning. |
Improving End-to-End Bangla Speech Recognition with Semi-supervised Training (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to train speech recognition systems require large annotated corpus. |
| Approach: | They propose a semi-supervised training approach that exploits large unpaired audio and text data to improve the performance of an automatic speech recognition system. |
| Outcome: | The proposed method reduces the WER of the system from 37% to 31.9%. |
Preparation of Bangla Speech Corpus from Publicly Available Audio & Text (2020.lrec-1)
Copied to clipboard
Shafayat Ahmed, Nafis Sadeq, Sudipta Saha Shubha, Md. Nahidul Islam, Muhammad Abdullah Adnan, Mohammad Zuberul Islam
| Challenge: | Automated speech recognition systems require large annotated speech corpus for training. |
| Approach: | They propose to use publicly available Bangla audiobooks and TV news recordings as input to prepare a large speech corpus with reasonable confidence. |
| Outcome: | The proposed algorithm outperforms the existing speech corpus and the existing corpus with speaker diarization and gender detection. |
Customizing Grapheme-to-Phoneme System for Non-Trivial Transcription Problems in Bangla Language (N19-1)
Copied to clipboard
Sudipta Saha Shubha, Nafis Sadeq, Shafayat Ahmed, Md. Nahidul Islam, Muhammad Abdullah Adnan, Md. Yasin Ali Khan, Mohammad Zuberul Islam
| Challenge: | Existing methods for Grapheme to phoneme conversion in Bangla language are mostly rule-based. |
| Approach: | They propose to use a lexicon to train a robust Grapheme to phoneme conversion system in Bangla language. |
| Outcome: | The proposed method outperforms other state-of-the-art approaches for G2P conversion in Bangla language. |