Papers by Md Sourove

2 papers
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)

Copied to clipboard

Challenge: low-resource languages like Bangla are limited by the lack of datasets.
Approach: They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla.
Outcome: The proposed datasets show that the proposed methods can enrich romanized Bangla.
BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla (2025.findings-naacl)

Copied to clipboard

Challenge: Existing work on monolingual or binary hate classification in Bangla has not addressed the challenge of multi-label hate speech classification in underrepresented languages.
Approach: They propose a multi-label transliterated Bangla hate speech dataset that translates or transliterates under-resourced text to higher-resource text before classifying the hate group(s).
Outcome: The proposed approach outperforms other methods in the zero-shot setting while achieving state-of-the-art performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations