Challenge: Identifying commercial posts in resource-constrained languages remains a challenge for automatic text classification tasks.
Approach: They propose a dataset for Bengali social media posts classified as commercial and noncommercial . they include an annotation guideline to aid future dataset creation in resource-constrained languages .
Outcome: The proposed dataset is based on an annotation guideline for future dataset creation in resource-constrained languages.

Similar Papers

Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)

Copied to clipboard

Challenge: resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task.
Approach: They present a survey of a multimodal dataset with different modalities according to the applications.
Outcome: The proposed datasets are available online and discuss the new frontier and motivate future researches.
BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification (2023.emnlp-main)

Copied to clipboard

Challenge: a number of studies have tried to detect and control the spread of such abusive memes on social media platforms.
Approach: They build a Bengali meme dataset to test models for abusive memes . they find that multimodal models that use both textual and visual information outperform unimodal models .
Outcome: The proposed model outperforms unimodal models in a Bengali meme dataset.
BD-SHS: A Benchmark Dataset for Learning to Detect Online Bangla Hate Speech in Different Social Contexts (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms and online streaming services have spawned a new breed of Hate Speech (HS) due to the massive amount of user-generated content, modern machine learning techniques are feasible and cost-effective to tackle this problem.
Approach: They propose to use a large manually labeled Bangla HS dataset to train generalizable models.
Outcome: The proposed dataset includes more than 50,200 offensive comments crawled from online social networking sites and is at least 60% larger than existing Bangla HS datasets.
MemoSen: A Multimodal Dataset for Sentiment Analysis of Memes (2022.lrec-1)

Copied to clipboard

Challenge: Recent studies on sentiment analysis of memes have focused on English, but there is a significant barrier to performing multimodal sentiment analysis research in resource-constrained languages like Bengali.
Approach: They propose to use a Bengali dataset to perform multimodal sentiment analysis in low resource languages.
Outcome: The proposed dataset for Bengali contains 4417 memes with three annotated labels positive, negative, and neutral.
MDS: A Fine-Grained Dataset for Multi-Modal Dialogue Summarization (2024.lrec-main)

Copied to clipboard

Challenge: Summarizing the dialogue into a short message has drawn much attention due to the explosion of various dialogue scenes.
Approach: They develop a multi-modal dialogue summarization dataset to enhance the variety of data available for this research area.
Outcome: The proposed dataset provides a demanding testbed for multi-modal dialogue summarization.
EmoNoBa: A Dataset for Analyzing Fine-Grained Emotions on Noisy Bangla Texts (2022.aacl-short)

Copied to clipboard

Challenge: EmoNoBa is a dataset for fine-grained emotion detection on Bangla text . it is based on 22698 comments from social media sites on 12 domains .
Approach: They propose a manually annotated dataset of 22,698 Bangla comments from social media sites on 12 different domains to use for fine-grained emotion detection.
Outcome: The proposed dataset of 22,698 public comments on 12 domains shows that hand-crafted features perform better than neural networks and pre-trained language models.
BanNERD: A Benchmark Dataset and Context-Driven Approach for Bangla Named Entity Recognition (2025.findings-naacl)

Copied to clipboard

Challenge: In a cross-dataset evaluation, models trained on BanNERD consistently outperformed those trained on four existing Bangla NER datasets.
Approach: They propose to use Bangla as a language to create the most extensive human-annotated and validated Bangla NLP dataset.
Outcome: The proposed method outperforms existing methods on Bangla NER datasets and performs competitively on English datasets.
MM-AVS: A Full-Scale Dataset for Multi-modal Summarization (2021.naacl-main)

Copied to clipboard

Challenge: Multimodal summarization materials lacking a holistic organization by integrating resources from various modalities.
Approach: They propose a multimodal article and video summarization dataset that integrates resources from different modalities.
Outcome: The proposed dataset validates the important assistance role of external information for multimodal summarization.
Can LLMs be Literary Companions?: Analysing LLMs on Bengali Figures of Speech Identification (2025.emnlp-main)

Copied to clipboard

Challenge: despite Bengali being among the most spoken languages, the NLP efforts on it remain limited.
Approach: They present a dataset that includes Bengali figures of speech on six poets . they deploy state-of-the-art Large Language Models to the dataset and fine-tune the best models .
Outcome: The proposed dataset reveals that two open-source LLMs perform better than others in Bengali . the framework can be reproduced for English and other low-resource languages .
Hate Speech and Offensive Language Detection in Bengali (2022.aacl-main)

Copied to clipboard

Challenge: Existing research on hate speech detection in English does not cover low-resource languages like Bengali.
Approach: They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets.
Outcome: The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations