Challenge: Bangla is the sixth most spoken language worldwide and the second Indo-Aryan language after Hindi.
Approach: They propose an annotated sentiment analysis dataset made of informally written Bangla texts.
Outcome: The proposed dataset is compared with neural networks and pretrained models . it shows that hand-crafted lexical features provide superior performance than neural networks .

Similar Papers

BanglaBook: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews (2023.findings-acl)

Copied to clipboard

Challenge: Existing literature on Bangla Sentiment Analysis (SA) has limited data and cross-domain adaptability.
Approach: They present a large-scale dataset of Bangla book reviews with 158,065 samples . they employ a range of machine learning models to establish baselines including SVM, LSTM, and Bangla-BERT.
Outcome: The proposed model improves performance over models that rely on manual features.
EmoNoBa: A Dataset for Analyzing Fine-Grained Emotions on Noisy Bangla Texts (2022.aacl-short)

Copied to clipboard

Challenge: EmoNoBa is a dataset for fine-grained emotion detection on Bangla text . it is based on 22698 comments from social media sites on 12 domains .
Approach: They propose a manually annotated dataset of 22,698 Bangla comments from social media sites on 12 different domains to use for fine-grained emotion detection.
Outcome: The proposed dataset of 22,698 public comments on 12 domains shows that hand-crafted features perform better than neural networks and pre-trained language models.
BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation (2026.acl-long)

Copied to clipboard

Challenge: Existing studies in Bangla focus on hate classification while overlooking interpretability.
Approach: They propose to create a dataset with human-annotated labels for banla that contains 19,203 YouTube comments spanning April 2024–June 2025.
Outcome: The proposed dataset outperforms existing datasets on open and closed-source LLMs on interpretability and better understanding of hate speech in linguistically rich yet under-resourced languages.
BD-SHS: A Benchmark Dataset for Learning to Detect Online Bangla Hate Speech in Different Social Contexts (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms and online streaming services have spawned a new breed of Hate Speech (HS) due to the massive amount of user-generated content, modern machine learning techniques are feasible and cost-effective to tackle this problem.
Approach: They propose to use a large manually labeled Bangla HS dataset to train generalizable models.
Outcome: The proposed dataset includes more than 50,200 offensive comments crawled from online social networking sites and is at least 60% larger than existing Bangla HS datasets.
User Guide for KOTE: Korean Online That-gul Emotions Dataset (2024.lrec-main)

Copied to clipboard

Challenge: sentiment analysis is used to identify emotional aspects of texts but is limited by its small size and limited range of emotions.
Approach: They propose a Korean sentiment analysis corpus that is limited by its small size and narrow range of emotions . they propose to fine-tune the KOTE dataset and analyze the results for social discrimination .
Outcome: The proposed dataset includes 50,000 Korean online comments, each manually labeled for 43 emotions and NO EMOTION.
BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla (2024.findings-emnlp)

Copied to clipboard

Challenge: low-resource languages like Bangla are limited by the lack of datasets.
Approach: They propose a large-scale transliteration dataset and a pre-training corpus on romanized Bangla.
Outcome: The proposed datasets show that the proposed methods can enrich romanized Bangla.
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla (2025.findings-acl)

Copied to clipboard

Challenge: ***BanStereoSet*** is a dataset designed to evaluate stereotypical social biases in multilingual LLMs for the Bangla language.
Approach: They propose to localize the content from StereoSet, IndiBias, and kamruzzaman-etal's datasets to capture biases prevalent within the Bangla language.
Outcome: The proposed dataset consists of 1,194 sentences spanning 9 categories of bias: race, profession, gender, ageism, beauty, beauty in profession, region, caste, and religion.
BanFakeNews: A Dataset for Detecting Fake News in Bangla (2020.lrec-1)

Copied to clipboard

Challenge: Impact of fake news is creating havoc worldwide.
Approach: They propose an annotated dataset of 50K news that can be used for building automated fake news detection systems for a low resource language like Bangla.
Outcome: The proposed system can be built with state-of-the-art NLP techniques for a low resource language like Bangla.
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment.
Approach: They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions.
Outcome: The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages.
The Design and Construction of a Chinese Sarcasm Dataset (2020.lrec-1)

Copied to clipboard

Challenge: Existing sarcasm datasets are limited to English and Chinese . sarcasm is a multi-layered semi-conscious language phenomenon .
Approach: They propose to build a high-quality Chinese sarcasm dataset using user comments . they use manual annotated sarkastic texts and non-sarcastic texts to train sarcasm classifier .
Outcome: The proposed dataset contains 2,486 manual annotated sarcastic texts and 89,296 non-sarcatic texts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations