SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Bangla is the sixth most spoken language worldwide and the second Indo-Aryan language after Hindi. |
| Approach: | They propose an annotated sentiment analysis dataset made of informally written Bangla texts. |
| Outcome: | The proposed dataset is compared with neural networks and pretrained models . it shows that hand-crafted lexical features provide superior performance than neural networks . |
Similar Papers
BanglaBook: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing literature on Bangla Sentiment Analysis (SA) has limited data and cross-domain adaptability. |
| Approach: | They present a large-scale dataset of Bangla book reviews with 158,065 samples . they employ a range of machine learning models to establish baselines including SVM, LSTM, and Bangla-BERT. |
| Outcome: | The proposed model improves performance over models that rely on manual features. |
EmoNoBa: A Dataset for Analyzing Fine-Grained Emotions on Noisy Bangla Texts (2022.aacl-short)
Copied to clipboard
| Challenge: | EmoNoBa is a dataset for fine-grained emotion detection on Bangla text . it is based on 22698 comments from social media sites on 12 domains . |
| Approach: | They propose a manually annotated dataset of 22,698 Bangla comments from social media sites on 12 different domains to use for fine-grained emotion detection. |
| Outcome: | The proposed dataset of 22,698 public comments on 12 domains shows that hand-crafted features perform better than neural networks and pre-trained language models. |
BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation (2026.acl-long)
Copied to clipboard
Faisal Hossain Raquib, Akm Moshiur Rahman Mazumder, Md Fahim, Md Tahmid Hasan Fuad, Md Farhan Ishmam, Faria Sultana, M Ashraful Amin, Amin Ahsan Ali, Akmmahbubur Rahman
| Challenge: | Existing studies in Bangla focus on hate classification while overlooking interpretability. |
| Approach: | They propose to create a dataset with human-annotated labels for banla that contains 19,203 YouTube comments spanning April 2024–June 2025. |
| Outcome: | The proposed dataset outperforms existing datasets on open and closed-source LLMs on interpretability and better understanding of hate speech in linguistically rich yet under-resourced languages. |
BD-SHS: A Benchmark Dataset for Learning to Detect Online Bangla Hate Speech in Different Social Contexts (2022.lrec-1)
Copied to clipboard
Nauros Romim, Mosahed Ahmed, Md Saiful Islam, Arnab Sen Sharma, Hriteshwar Talukder, Mohammad Ruhul Amin
| Challenge: | Social media platforms and online streaming services have spawned a new breed of Hate Speech (HS) due to the massive amount of user-generated content, modern machine learning techniques are feasible and cost-effective to tackle this problem. |
| Approach: | They propose to use a large manually labeled Bangla HS dataset to train generalizable models. |
| Outcome: | The proposed dataset includes more than 50,200 offensive comments crawled from online social networking sites and is at least 60% larger than existing Bangla HS datasets. |
User Guide for KOTE: Korean Online That-gul Emotions Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | sentiment analysis is used to identify emotional aspects of texts but is limited by its small size and limited range of emotions. |
| Approach: | They propose a Korean sentiment analysis corpus that is limited by its small size and narrow range of emotions . they propose to fine-tune the KOTE dataset and analyze the results for social discrimination . |
| Outcome: | The proposed dataset includes 50,000 Korean online comments, each manually labeled for 43 emotions and NO EMOTION. |
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)
Copied to clipboard
| Challenge: | low-resource languages like Bangla are limited by the lack of datasets. |
| Approach: | They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla. |
| Outcome: | The proposed datasets show that the proposed methods can enrich romanized Bangla. |
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla (2025.findings-acl)
Copied to clipboard
| Challenge: | ***BanStereoSet*** is a dataset designed to evaluate stereotypical social biases in multilingual LLMs for the Bangla language. |
| Approach: | They propose to localize the content from StereoSet, IndiBias, and kamruzzaman-etal's datasets to capture biases prevalent within the Bangla language. |
| Outcome: | The proposed dataset consists of 1,194 sentences spanning 9 categories of bias: race, profession, gender, ageism, beauty, beauty in profession, region, caste, and religion. |
BanFakeNews: A Dataset for Detecting Fake News in Bangla (2020.lrec-1)
Copied to clipboard
| Challenge: | Impact of fake news is creating havoc worldwide. |
| Approach: | They propose an annotated dataset of 50K news that can be used for building automated fake news detection systems for a low resource language like Bangla. |
| Outcome: | The proposed system can be built with state-of-the-art NLP techniques for a low resource language like Bangla. |
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment. |
| Approach: | They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions. |
| Outcome: | The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages. |
The Design and Construction of a Chinese Sarcasm Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing sarcasm datasets are limited to English and Chinese . sarcasm is a multi-layered semi-conscious language phenomenon . |
| Approach: | They propose to build a high-quality Chinese sarcasm dataset using user comments . they use manual annotated sarkastic texts and non-sarcastic texts to train sarcasm classifier . |
| Outcome: | The proposed dataset contains 2,486 manual annotated sarcastic texts and 89,296 non-sarcatic texts. |