Challenge: Recent studies have shown that Large Language Models (LLMs) neglect cultural diversity and key aspects like emotion and contextual knowledge hidden in the visual modalities.
Approach: They propose a framework for misogynous meme identification using a multimodal multimodal prompting principle and a CLIP-based classifier.
Outcome: The proposed framework performs well on the SemEval-2022 task 5 dataset, and is generalizable across different datasets.

Similar Papers

Unintended Bias Detection and Mitigation in Misogynous Memes (2024.eacl-long)

Copied to clipboard

Challenge: Existing models that detect misogyny are not able to detect unintended biases in memes, perpetuating harmful stereotypes and reinforcing negative attitudes.
Approach: They propose to measure and mitigate unintentional bias in misogynous memes detection models by using a contextualized scene graph-based multimodal network (CTXSGMNet) they also evaluate their generalizability by evaluating their performance on a few benchmark meme datasets.
Outcome: The proposed model achieves state-of-the-art performance on the SemEval-2022 Task 5 (MAMI task) dataset, showcasing its promising performance in terms of Equity of Odds and F1 score.
MemeWeaver: Inter-Meme Graph Reasoning for Sexism and Misogyny Detection (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods to detect hate speech on social media are limited by heuristic graph construction, shallow modality fusion, and instance-level reasoning.
Approach: They propose a multimodal framework for detecting sexism and misogyny using a graph reasoning mechanism that can be used to train multiple visual-textual fusion strategies.
Outcome: The proposed framework outperforms state-of-the-art methods on MAMI and EXIST benchmarks while achieving faster training convergence.
Exploring Chain-of-Thought for Multi-modal Metaphor Detection (2024.acl-long)

Copied to clipboard

Challenge: Metaphors are commonly found in advertising and internet memes, but lack of high-quality textual data is a challenge for language models . a new framework for multi-modal metaphor detection is being developed to address these challenges .
Approach: They propose a framework that extracts and integrates knowledge from Large Language Models into smaller ones to improve model performance.
Outcome: The proposed framework outperforms existing models on the MET-MEME dataset.
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification (2024.emnlp-main)

Copied to clipboard

Challenge: a novel dataset of text-embedded images associated with the LGBTQ+ Pride movement is presented in this paper . a new framework for analyzing text-based images is proposed to address this challenge .
Approach: They propose a new dataset for machine learning that includes hate, targets of hate, stance, humor and a framework for efficient downstream learning while preserving the knowledge of the pre-trained CLIP model.
Outcome: The proposed framework achieves superior performance on two real-world datasets.
Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for harmful meme detection ignore in-depth cognition of meme text and image . authors propose a framework for learning reasonable thoughts from LLMs for better multimodal fusion .
Approach: They propose to use large language models to learn reasonable thoughts from LLMs for better multimodal fusion and lightweight fine-tuning.
Outcome: The proposed approach achieves superior performance than state-of-the-art methods on the harmful meme detection task.
BANMIME : Misogyny Detection with Metaphor Explanation on Bangla Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have explored hate speech and general meme classification, but the nuanced identification of misogyny in Bangla memes remains underexplored.
Approach: They propose a Bangla misogynistic meme dataset that includes misos, humor, metaphors and detailed human-written explanations.
Outcome: The proposed dataset is the first comprehensive dataset of misogynistic Bangla memes . it includes misos, humor categories, metaphor localization, and detailed human-written explanations based on 2,000 culturally grounded samples .
Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detection (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for detecting hateful memes rely on extensive training.
Approach: They propose a method that integrates evolution attribute and in-context information of memes into large multimodal models via Chain-of-Evolution (CoE) prompting.
Outcome: The proposed method improves existing methods on public datasets and can be used as interpretive tool to promote understanding of evolution of memes.
Revealing the Seen, Imagining the Beyond: A Survey of Image-Grounded Chain-of-Thought Reasoning in Multimodal LLMs (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Multimodal Large Language Models (MLLMs) have shifted visual reasoning from tool-calling to end-to-end perceptionreasoning.
Approach: They synthesize the emerging paradigm of Image-Grounded Chain-of-Thought (IG-CoT) they propose a method-centric taxonomy covering prompting, supervised fine-tuning, and reinforcement learning .
Outcome: The proposed model is based on a method-centric taxonomy and benchmarks.
Chain-of-Thought Embeddings for Stance Detection on Social Media (2023.findings-emnlp)

Copied to clipboard

Challenge: Stance detection on social media platforms like Twitter is challenging for Large Language Models (LLMs), as emerging slang and colloquial language in online conversations often contain deeply implicit stance labels.
Approach: They propose to embed COT reasonings into a traditional RoBERTa-based stance detection pipeline by embedding COT stance reasonings and integrating them into slang-based models.
Outcome: The proposed model achieves SOTA performance on multiple stance detection datasets collected from social media.
A Multimodal Framework to Detect Target Aware Aggression in Memes (2024.eacl-long)

Copied to clipboard

Challenge: Recent research on memes’ detrimental facets is skewed towards high-resource languages, such as Bengali.
Approach: They propose a dataset MIMOSA that annotates annotated memes across five aggression target categories in Bengali and propose 'Multimodal Attentive Fusion' to detect aggression targets.
Outcome: The proposed method outperforms state-of-the-art methods in Bengali and in low-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations