MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation approaches focus on mLLMs’ detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts. |
| Approach: | They propose an agent-based arena-style evaluation framework that provides context-aware and unbiased assessment for mLLMs’ understanding of multimodal harmfulness. |
| Outcome: | The proposed framework reduces evaluation biases of judge agents and provides unbiased comparisons of mLLMs’ abilities to interpret multimodal harmfulness. |
Similar Papers
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness (2025.acl-long)
Copied to clipboard
| Challenge: | Existing models that assess mLLMs on harmful meme understanding are inaccurate and lack accuracy. |
| Approach: | They propose a framework that adaptively probes the reasoning capabilities of mLLMs . their framework systematically reveals the varying performance of different target mllms a . |
| Outcome: | The proposed framework systematically reveals the performance of different target mLLMs. |
Beneath the Surface: Unveiling Harmful Memes with Multimodal Reasoning Distilled from Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for harmful meme detection ignore in-depth cognition of meme text and image . authors propose a framework for learning reasonable thoughts from LLMs for better multimodal fusion . |
| Approach: | They propose to use large language models to learn reasonable thoughts from LLMs for better multimodal fusion and lightweight fine-tuning. |
| Outcome: | The proposed approach achieves superior performance than state-of-the-art methods on the harmful meme detection task. |
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria (2025.naacl-long)
Copied to clipboard
Wentao Ge, Shunian Chen, Hardy Chen, Nuo Chen, Junying Chen, Zhihong Chen, Wenya Xie, Shuo Yan, ChenghaoZhu ChenghaoZhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang
| Challenge: | Existing evaluation methodologies for multimodal large language models are limited in evaluating objective queries without considering real-world user experiences. |
| Approach: | They propose to evaluate multimodal large language models with per-sample criteria using potent MLLM as the judge. |
| Outcome: | The proposed evaluation paradigm shows that it can be used to evaluate multimodal large language models with per-sample criteria. |
MM-SOC: Benchmarking Multimodal Large Language Models in Social Media Platforms (2024.findings-acl)
Copied to clipboard
| Challenge: | Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces. |
| Approach: | They propose a benchmark to evaluate MLLMs' understanding of multimodal social media content and a large-scale YouTube tagging dataset to evaluate their performance. |
| Outcome: | The proposed model performs better in a zero-shot setting, suggesting potential improvements. |
MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization (2023.acl-long)
Copied to clipboard
| Challenge: | Besides digital archiving of memes and their metadata, there is no efficient way to deduce a meme’s context dynamically. |
| Approach: | They propose a task to mine the context that succinctly explains the background of a meme and a related document to capture cross-modal semantic dependencies between the meme and the context. |
| Outcome: | The proposed dataset outperforms existing systems and shows that it can capture cross-modal semantic dependencies between the meme and the context. |
MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets (2021.findings-emnlp)
Copied to clipboard
Shraman Pramanick, Shivam Sharma, Dimitar Dimitrov, Md. Shad Akhtar, Preslav Nakov, Tanmoy Chakraborty
| Challenge: | a growing number of harmful memes are being used for trolling, cyberbullying and abuse . a new approach to detect harmful meme images and texts is emerging . |
| Approach: | They propose a multimodal deep neural network that detects harmful memes . they extend the recently released HarMeme dataset with additional memes and a new topic . |
| Outcome: | The proposed framework outperforms rival methods in detecting harmful memes and their target social entities. |
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies on content moderation of toxic memes focus on text-based content . current research neglects the widespread influence of multimodal content like memes . |
| Approach: | They propose a framework leveraging Large Language Models and Visual Language Model (VLMs) for meme intervention. |
| Outcome: | The proposed framework enables users to generate relevant and effective responses to toxic memes. |
Towards Low-Resource Harmful Meme Detection with LMM Agents (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for harmful meme detection are limited due to the dynamic nature of memes . eliciting knowledge-revising behavior within the LMM agent is a key factor in achieving this goal . |
| Approach: | They propose an agency-driven framework for low-resource harmful meme detection . they use annotated memes to leverage label information as auxiliary signals for model . |
| Outcome: | The proposed framework achieves superior performance than state-of-the-art methods on the low-resource harmful meme detection task. |
MemeQA: Holistic Evaluation for Meme Understanding (2025.acl-long)
Copied to clipboard
Khoi P. N. Nguyen, Terrence Li, Derek Lou Zhou, Gabriel Xiong, Pranav Balu, Nandhan Alahari, Alan Huang, Tanush Chauhan, Harshavardhan Bala, Emre Guzelordu, Affan Kashfi, Aaron Xu, Suyesh Shrestha, Megan Vu, Jerry Wang, Vincent Ng
| Challenge: | Existing benchmarks for meme understanding only concern narrow aspects of meme semantics. |
| Approach: | They propose to use multiple-choice questions to evaluate meme comprehension . they use a dataset of over 9,000 multiple-question questions to assess meme comprehension. |
| Outcome: | The proposed model outperforms existing models on meme comprehension . the model makes many errors on memes where proper understanding requires going beyond sentiment . |
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have focused on harms of memes in closed environments, such as hate speech and cyber-bullying. |
| Approach: | They propose a multimodal question-answering framework that solicits accurate responses to structured questions while providing coherent explanations. |
| Outcome: | The proposed framework outperforms existing frameworks in predicting answer prediction accuracy and text generation lead over a baseline. |