Challenge: Existing models that detect misogyny are not able to detect unintended biases in memes, perpetuating harmful stereotypes and reinforcing negative attitudes.
Approach: They propose to measure and mitigate unintentional bias in misogynous memes detection models by using a contextualized scene graph-based multimodal network (CTXSGMNet) they also evaluate their generalizability by evaluating their performance on a few benchmark meme datasets.
Outcome: The proposed model achieves state-of-the-art performance on the SemEval-2022 Task 5 (MAMI task) dataset, showcasing its promising performance in terms of Equity of Odds and F1 score.

Similar Papers

MemeWeaver: Inter-Meme Graph Reasoning for Sexism and Misogyny Detection (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods to detect hate speech on social media are limited by heuristic graph construction, shallow modality fusion, and instance-level reasoning.
Approach: They propose a multimodal framework for detecting sexism and misogyny using a graph reasoning mechanism that can be used to train multiple visual-textual fusion strategies.
Outcome: The proposed framework outperforms state-of-the-art methods on MAMI and EXIST benchmarks while achieving faster training convergence.
From Laughter to Inequality: Annotated Dataset for Misogyny Detection in Tamil and Malayalam Memes (2024.lrec-main)

Copied to clipboard

Challenge: a new form of memes has emerged to combat misogyny and harmful stereotypes . authors present a dataset to analyze online misogamy in Tamil and Malayalam communities .
Approach: They propose to create an annotated dataset with detailed annotation guidelines to analyze online misogyny within Tamil and Malayalam-speaking communities.
Outcome: The proposed dataset reveals the world of gender bias and stereotypes in Tamil and Malayalam-speaking communities.
M3Hop-CoT: Misogynous Meme Identification with Multimodal Multi-hop Chain-of-Thought (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that Large Language Models (LLMs) neglect cultural diversity and key aspects like emotion and contextual knowledge hidden in the visual modalities.
Approach: They propose a framework for misogynous meme identification using a multimodal multimodal prompting principle and a CLIP-based classifier.
Outcome: The proposed framework performs well on the SemEval-2022 task 5 dataset, and is generalizable across different datasets.
Seeing Through VisualBERT: A Causal Adventure on Memetic Landscapes (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models for detecting offensive memes lack transparency and are often unreliability in safety-critical applications.
Approach: They propose a framework that uses a Structural Causal Model to predict the class of an input meme based on meme input and causal concepts, allowing for transparent interpretation.
Outcome: The proposed framework is able to predict class of an input meme based on meme input and causal concepts, allowing for transparent interpretation.
MemeDetoxNet: Balancing Toxicity Reduction and Context Preservation (2025.findings-acl)

Copied to clipboard

Challenge: Toxic memes spread harmful and offensive content and pose a significant challenge in online environments.
Approach: They propose a framework to mitigate toxicity in toxic memes by leveraging a set of pre-trained models that can interpret the visual and textual components of memes.
Outcome: The proposed framework reduces toxicity on publicly available meme datasets by 10-20% compared to the previous methods.
DISARM: Detecting the Victims Targeted by Harmful Memes (2022.findings-naacl)

Copied to clipboard

Challenge: DISARM is a framework that uses named-entity recognition and person identification to detect all entities a meme is referring to and then incorporates a novel contextualized deep neural network to classify whether the meme intends to harm these entities.
Approach: They propose a framework that uses named-entity recognition and person identification to detect all entities a meme is referring to and incorporates a novel contextualized deep neural network to classify whether the meme intends to harm them.
Outcome: The proposed framework outperforms 10 unimodal and multimodal systems and reduces error rate of harmful target identification by 9 % absolute over baseline systems.
MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization (2023.acl-long)

Copied to clipboard

Challenge: Besides digital archiving of memes and their metadata, there is no efficient way to deduce a meme’s context dynamically.
Approach: They propose a task to mine the context that succinctly explains the background of a meme and a related document to capture cross-modal semantic dependencies between the meme and the context.
Outcome: The proposed dataset outperforms existing systems and shows that it can capture cross-modal semantic dependencies between the meme and the context.
MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets (2021.findings-emnlp)

Copied to clipboard

Challenge: a growing number of harmful memes are being used for trolling, cyberbullying and abuse . a new approach to detect harmful meme images and texts is emerging .
Approach: They propose a multimodal deep neural network that detects harmful memes . they extend the recently released HarMeme dataset with additional memes and a new topic .
Outcome: The proposed framework outperforms rival methods in detecting harmful memes and their target social entities.
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for label detection and explanation generation have been limited in understanding complex issues . identifying propaganda and hate in memes is essential for combating misinformation and minimizing harm .
Approach: They propose an explanation-enhanced dataset for propaganda memes in Arabic and hateful memes on English to solve these tasks.
Outcome: The proposed model outperforms the current state-of-the-art in label detection and explanation generation.
A Context-Aware Contrastive Learning Framework for Hateful Meme Detection and Segmentation (2025.findings-naacl)

Copied to clipboard

Challenge: Empirical experiments show HateSieve surpasses existing LMMs in performance with fewer trainable parameters .
Approach: They propose a framework to enhance detection and segmentation of hateful elements in memes by creating a triplet dataset and an Image-Text Alignment module.
Outcome: HateSieve features a new framework that creates semantically correlated memes and generates contextual embeddings for accurate meme segmentation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations