Challenge: Empirical experiments show HateSieve surpasses existing LMMs in performance with fewer trainable parameters .
Approach: They propose a framework to enhance detection and segmentation of hateful elements in memes by creating a triplet dataset and an Image-Text Alignment module.
Outcome: HateSieve features a new framework that creates semantically correlated memes and generates contextual embeddings for accurate meme segmentation.

Similar Papers

Improving Hateful Meme Detection through Retrieval-Guided Contrastive Learning (2024.acl-long)

Copied to clipboard

Challenge: Existing systems for detecting hateful memes lack sensitivity to subtle differences in memes that are vital for correct hatefulness classification.
Approach: They propose to construct a hatefulness-aware embedding space through retrieval-guided contrastive training to identify hatefulness based on data unseen in training.
Outcome: The proposed system outperforms existing models on the HatefulMemes dataset with an AUROC of 87.0 and improves contextual understanding across domains.
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection (2025.emnlp-main)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) have shown promise in hateful meme detection, but they face limitations like sub-optimal performance and limited out-of-domain generalization capabilities.
Approach: They propose a robust adaptation framework for hateful meme detection that enhances in-domain accuracy and cross-domain generalization while preserving the general vision-language capabilities of LMMs.
Outcome: The proposed framework outperforms larger agentic systems in detecting hateful memes under adversarial attacks while maintaining the general vision-language capabilities of LMMs.
Caption Enriched Samples for Improving Hateful Memes Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy.
Approach: They propose to use an off-the-shelf caption generator to capture the first image and overlayed text.
Outcome: The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text .
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for label detection and explanation generation have been limited in understanding complex issues . identifying propaganda and hate in memes is essential for combating misinformation and minimizing harm .
Approach: They propose an explanation-enhanced dataset for propaganda memes in Arabic and hateful memes on English to solve these tasks.
Outcome: The proposed model outperforms the current state-of-the-art in label detection and explanation generation.
Deciphering Hate: Identifying Hateful Memes and Their Targets (2024.acl-long)

Copied to clipboard

Challenge: a growing body of research has focused on the negative aspects of memes in high-resource languages like Bengali . a new dataset for Bengali hateful memes is designed to detect their targeted entities .
Approach: They propose a multimodal dataset that analyzes the modality of memes and compares them with other datasets.
Outcome: The proposed dataset outperforms state-of-the-art datasets on Bengali hateful memes . the proposed dataset is generalizable on other low-resource hateful memes datasets compared with baselines based on the proposed model .
Text or Image? What is More Important in Cross-Domain Generalization Capabilities of Hate Meme Detection Models? (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies show that only the textual component of hateful memes enables the multimodal classifier to generalize across domains while the image component proves highly sensitive to a specific training dataset.
Approach: They propose to use only the textual component of hateful memes to generalize across different domains while the image component is highly sensitive to a specific training dataset.
Outcome: The proposed model performs similarly to hate-meme classifiers in a zero-shot setting, while the introduction of meme’s image captions worsens performance by an average F1 of 0.02.
Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection (2024.eacl-srw)

Copied to clipboard

Challenge: Existing approaches to multimodal hateful content detection focus on detecting hate speech from text-based content, but they fail to address modality-specific features.
Approach: They propose a context-aware attention framework for multimodal hateful content detection that integrates an attention layer to meaningfully align the visual and textual features.
Outcome: The proposed framework achieves F1-scores of 69.7% and 70.3% on two hateful meme datasets and shows 2.5% and 3.2% performance improvement over the state-of-the-art systems.
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification (2024.emnlp-main)

Copied to clipboard

Challenge: a novel dataset of text-embedded images associated with the LGBTQ+ Pride movement is presented in this paper . a new framework for analyzing text-based images is proposed to address this challenge .
Approach: They propose a new dataset for machine learning that includes hate, targets of hate, stance, humor and a framework for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model.
Outcome: The proposed framework achieves superior performance on two real-world datasets.
Deciphering Implicit Hate: Evaluating Automated Detection Algorithms for Multimodal Hate (2021.findings-acl)

Copied to clipboard

Challenge: Imlicit hate content has unusual syntax, polysemic words, and fewer markers of prejudice, e.g., slurs . multimodal content is harder to detect than unimodal content, such as memes .
Approach: They evaluate the role of semantic and multimodal context for detecting implicit and explicit hate . they find that all models perform better on content with full annotator agreement .
Outcome: The proposed model outperforms other models on implicit and explicit hate detection tasks because of its lower propensity towards false positives.
MUTE: A Multimodal Dataset for Detecting Hateful Memes (2022.aacl-srw)

Copied to clipboard

Challenge: social media has enabled information propagation at unprecedented rate, but also generated malign content, such as hateful memes . a multimodal hate speech dataset is used to study the impact of hateful content on society . current studies focus on monolingual memes, but existing models cannot provide accurate inferences based on code-mixed captions a study on Bengali memes shows that joint evaluation of visual and textual features significantly improves the hateful data classification .
Approach: They propose to use a multimodal hate speech dataset to detect hateful memes . they use monolingual captions in English and Bengali to analyze the content .
Outcome: The proposed dataset shows that evaluation of visual and textual features significantly improves the hateful memes classification compared to unimodal evaluation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations