Challenge: social media has enabled information propagation at unprecedented rate, but also generated malign content, such as hateful memes . a multimodal hate speech dataset is used to study the impact of hateful content on society . current studies focus on monolingual memes, but existing models cannot provide accurate inferences based on code-mixed captions a study on Bengali memes shows that joint evaluation of visual and textual features significantly improves the hateful data classification .
Approach: They propose to use a multimodal hate speech dataset to detect hateful memes . they use monolingual captions in English and Bengali to analyze the content .
Outcome: The proposed dataset shows that evaluation of visual and textual features significantly improves the hateful memes classification compared to unimodal evaluation.

Similar Papers

Deciphering Hate: Identifying Hateful Memes and Their Targets (2024.acl-long)

Copied to clipboard

Challenge: a growing body of research has focused on the negative aspects of memes in high-resource languages like Bengali . a new dataset for Bengali hateful memes is designed to detect their targeted entities .
Approach: They propose a multimodal dataset that analyzes the modality of memes and compares them with other datasets.
Outcome: The proposed dataset outperforms state-of-the-art datasets on Bengali hateful memes . the proposed dataset is generalizable on other low-resource hateful memes datasets compared with baselines based on the proposed model .
MemoSen: A Multimodal Dataset for Sentiment Analysis of Memes (2022.lrec-1)

Copied to clipboard

Challenge: Recent studies on sentiment analysis of memes have focused on English, but there is a significant barrier to performing multimodal sentiment analysis research in resource-constrained languages like Bengali.
Approach: They propose to use a Bengali dataset to perform multimodal sentiment analysis in low resource languages.
Outcome: The proposed dataset for Bengali contains 4417 memes with three annotated labels positive, negative, and neutral.
A Multimodal Framework to Detect Target Aware Aggression in Memes (2024.eacl-long)

Copied to clipboard

Challenge: Recent research on memes’ detrimental facets is skewed towards high-resource languages, such as Bengali.
Approach: They propose a dataset MIMOSA that annotates annotated memes across five aggression target categories in Bengali and propose 'Multimodal Attentive Fusion' to detect aggression targets.
Outcome: The proposed method outperforms state-of-the-art methods in Bengali and in low-resource languages.
Caption Enriched Samples for Improving Hateful Memes Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy.
Approach: They propose to use an off-the-shelf caption generator to capture the first image and overlayed text.
Outcome: The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text .
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection (2024.eacl-srw)

Copied to clipboard

Challenge: Existing approaches to multimodal hateful content detection focus on detecting hate speech from text-based content, but they fail to address modality-specific features.
Approach: They propose a context-aware attention framework for multimodal hateful content detection that integrates an attention layer to meaningfully align the visual and textual features.
Outcome: The proposed framework achieves F1-scores of 69.7% and 70.3% on two hateful meme datasets and shows 2.5% and 3.2% performance improvement over the state-of-the-art systems.
A Context-Aware Contrastive Learning Framework for Hateful Meme Detection and Segmentation (2025.findings-naacl)

Copied to clipboard

Challenge: Empirical experiments show HateSieve surpasses existing LMMs in performance with fewer trainable parameters .
Approach: They propose a framework to enhance detection and segmentation of hateful elements in memes by creating a triplet dataset and an Image-Text Alignment module.
Outcome: HateSieve features a new framework that creates semantically correlated memes and generates contextual embeddings for accurate meme segmentation.
BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification (2023.emnlp-main)

Copied to clipboard

Challenge: a number of studies have tried to detect and control the spread of such abusive memes on social media platforms.
Approach: They build a Bengali meme dataset to test models for abusive memes . they find that multimodal models that use both textual and visual information outperform unimodal models .
Outcome: The proposed model outperforms unimodal models in a Bengali meme dataset.
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision–Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a new study shows that cultural background significantly affects multimodal hate speech moderation models . a limited dataset excludes multi-modal forms of hate and excludes non-English-speaking cultures . the lowest pairwise label agreement between the USA and India is due to cultural factors .
Approach: They use a multimodal and multilingual parallel hate speech dataset to examine cultural differences . they find that cultural background significantly affects multimodal hate speech annotation .
Outcome: The proposed dataset shows that cultural background significantly affects multimodal hate speech annotation.
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification (2024.emnlp-main)

Copied to clipboard

Challenge: a novel dataset of text-embedded images associated with the LGBTQ+ Pride movement is presented in this paper . a new framework for analyzing text-based images is proposed to address this challenge .
Approach: They propose a new dataset for machine learning that includes hate, targets of hate, stance, humor and a framework for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model.
Outcome: The proposed framework achieves superior performance on two real-world datasets.
Hate Speech and Offensive Language Detection in Bengali (2022.aacl-main)

Copied to clipboard

Challenge: Existing research on hate speech detection in English does not cover low-resource languages like Bengali.
Approach: They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets.
Outcome: The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations