Papers by Md Sourove
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)
Copied to clipboard
| Challenge: | low-resource languages like Bangla are limited by the lack of datasets. |
| Approach: | They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla. |
| Outcome: | The proposed datasets show that the proposed methods can enrich romanized Bangla. |
BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla (2025.findings-naacl)
Copied to clipboard
Fabiha Haider, Fariha Tanjim Shifat, Md Farhan Ishmam, Md Sakib Ul Rahman Sourove, Deeparghya Dutta Barua, Md Fahim, Md Farhad Alam Bhuiyan
| Challenge: | Existing work on monolingual or binary hate classification in Bangla has not addressed the challenge of multi-label hate speech classification in underrepresented languages. |
| Approach: | They propose a multi-label transliterated Bangla hate speech dataset that translates or transliterates under-resourced text to higher-resource text before classifying the hate group(s). |
| Outcome: | The proposed approach outperforms other methods in the zero-shot setting while achieving state-of-the-art performance. |