| Challenge: | social media has enabled information propagation at unprecedented rate, but also generated malign content, such as hateful memes . a multimodal hate speech dataset is used to study the impact of hateful content on society . current studies focus on monolingual memes, but existing models cannot provide accurate inferences based on code-mixed captions a study on Bengali memes shows that joint evaluation of visual and textual features significantly improves the hateful data classification . |
| Approach: | They propose to use a multimodal hate speech dataset to detect hateful memes . they use monolingual captions in English and Bengali to analyze the content . |
| Outcome: | The proposed dataset shows that evaluation of visual and textual features significantly improves the hateful memes classification compared to unimodal evaluation. |
Similar Papers
Deciphering Hate: Identifying Hateful Memes and Their Targets (2024.acl-long)
Copied to clipboard
| Challenge: | a growing body of research has focused on the negative aspects of memes in high-resource languages like Bengali . a new dataset for Bengali hateful memes is designed to detect their targeted entities . |
| Approach: | They propose a multimodal dataset that analyzes the modality of memes and compares them with other datasets. |
| Outcome: | The proposed dataset outperforms state-of-the-art datasets on Bengali hateful memes . the proposed dataset is generalizable on other low-resource hateful memes datasets compared with baselines based on the proposed model . |
MemoSen: A Multimodal Dataset for Sentiment Analysis of Memes (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies on sentiment analysis of memes have focused on English, but there is a significant barrier to performing multimodal sentiment analysis research in resource-constrained languages like Bengali. |
| Approach: | They propose to use a Bengali dataset to perform multimodal sentiment analysis in low resource languages. |
| Outcome: | The proposed dataset for Bengali contains 4417 memes with three annotated labels positive, negative, and neutral. |
A Multimodal Framework to Detect Target Aware Aggression in Memes (2024.eacl-long)
Copied to clipboard
| Challenge: | Recent research on memes’ detrimental facets is skewed towards high-resource languages, such as Bengali. |
| Approach: | They propose a dataset MIMOSA that annotates annotated memes across five aggression target categories in Bengali and propose 'Multimodal Attentive Fusion' to detect aggression targets. |
| Outcome: | The proposed method outperforms state-of-the-art methods in Bengali and in low-resource languages. |
Caption Enriched Samples for Improving Hateful Memes Detection (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy. |
| Approach: | They propose to use an off-the-shelf caption generator to capture the first image and overlayed text. |
| Outcome: | The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text . |
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection (2024.eacl-srw)
Copied to clipboard
| Challenge: | Existing approaches to multimodal hateful content detection focus on detecting hate speech from text-based content, but they fail to address modality-specific features. |
| Approach: | They propose a context-aware attention framework for multimodal hateful content detection that integrates an attention layer to meaningfully align the visual and textual features. |
| Outcome: | The proposed framework achieves F1-scores of 69.7% and 70.3% on two hateful meme datasets and shows 2.5% and 3.2% performance improvement over the state-of-the-art systems. |
A Context-Aware Contrastive Learning Framework for Hateful Meme Detection and Segmentation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Empirical experiments show HateSieve surpasses existing LMMs in performance with fewer trainable parameters . |
| Approach: | They propose a framework to enhance detection and segmentation of hateful elements in memes by creating a triplet dataset and an Image-Text Alignment module. |
| Outcome: | HateSieve features a new framework that creates semantically correlated memes and generates contextual embeddings for accurate meme segmentation. |
BanglaAbuseMeme: A Dataset for Bengali Abusive Meme Classification (2023.emnlp-main)
Copied to clipboard
| Challenge: | a number of studies have tried to detect and control the spread of such abusive memes on social media platforms. |
| Approach: | They build a Bengali meme dataset to test models for abusive memes . they find that multimodal models that use both textual and visual information outperform unimodal models . |
| Outcome: | The proposed model outperforms unimodal models in a Bengali meme dataset. |
Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision–Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | a new study shows that cultural background significantly affects multimodal hate speech moderation models . a limited dataset excludes multi-modal forms of hate and excludes non-English-speaking cultures . the lowest pairwise label agreement between the USA and India is due to cultural factors . |
| Approach: | They use a multimodal and multilingual parallel hate speech dataset to examine cultural differences . they find that cultural background significantly affects multimodal hate speech annotation . |
| Outcome: | The proposed dataset shows that cultural background significantly affects multimodal hate speech annotation. |
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification (2024.emnlp-main)
Copied to clipboard
| Challenge: | a novel dataset of text-embedded images associated with the LGBTQ+ Pride movement is presented in this paper . a new framework for analyzing text-based images is proposed to address this challenge . |
| Approach: | They propose a new dataset for machine learning that includes hate, targets of hate, stance, humor and a framework for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model. |
| Outcome: | The proposed framework achieves superior performance on two real-world datasets. |
Hate Speech and Offensive Language Detection in Bengali (2022.aacl-main)
Copied to clipboard
| Challenge: | Existing research on hate speech detection in English does not cover low-resource languages like Bengali. |
| Approach: | They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets. |
| Outcome: | The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better. |