MemeQA: Holistic Evaluation for Meme Understanding (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for meme understanding only concern narrow aspects of meme semantics.
Approach: They propose to use multiple-choice questions to evaluate meme comprehension . they use a dataset of over 9,000 multiple-question questions to assess meme comprehension.
Outcome: The proposed model outperforms existing models on meme comprehension . the model makes many errors on memes where proper understanding requires going beyond sentiment .

Similar Papers

MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have focused on harms of memes in closed environments, such as hate speech and cyber-bullying.
Approach: They propose a multimodal question-answering framework that solicits accurate responses to structured questions while providing coherent explanations.
Outcome: The proposed framework outperforms existing frameworks in predicting answer prediction accuracy and text generation lead over a baseline.
Computational Meme Understanding: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Computational Meme Understanding (CMU) is a collection of tasks involving the automated comprehension of memes.
Approach: They propose a comprehensive taxonomy for memes along three dimensions – forms, functions, and topics and introduce three key tasks for Computational Meme Understanding, namely classification, interpretation, and explanation.
Outcome: The proposed model is based on a taxonomy of memes along three dimensions and is compared to existing models and datasets.
MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization (2023.acl-long)

Copied to clipboard

Challenge: Besides digital archiving of memes and their metadata, there is no efficient way to deduce a meme’s context dynamically.
Approach: They propose a task to mine the context that succinctly explains the background of a meme and a related document to capture cross-modal semantic dependencies between the meme and the context.
Outcome: The proposed dataset outperforms existing systems and shows that it can capture cross-modal semantic dependencies between the meme and the context.
MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Current approaches focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overlooking a fundamental challenge: the same meme can express different intents depending on its conversational context.
Approach: They propose a benchmark to evaluate how large vision language models understand memes in their original context.
Outcome: The proposed benchmark evaluates how large vision language models understand meme intent in their original context.
MemeCap: A Dataset for Captioning and Interpreting Memes (2023.emnlp-main)

Copied to clipboard

Challenge: a new dataset aims to understand meme captioning tasks using visual metaphors . vision and language models are proving to be effective in image captioning and visual question answering tasks .
Approach: They present a dataset that contains 6.3K memes and 6.3k meme captions . they show that vision and language models still struggle with visual metaphors despite their advanced capabilities .
Outcome: The proposed dataset contains 6.3K memes along with the title of the post containing the meme, meme captions, literal image caption, and visual metaphors.
MemeInterpret: Towards an All-in-One Dataset for Meme Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing research has not explored meme captioning's decomposition into subtasks or its connections to other CMU tasks.
Approach: a new meme corpus is built upon the Facebook Hateful Memes dataset . it contains meme captions, corresponding surface messages and relevant background knowledge .
Outcome: a new corpus of meme captions and surface messages unifies three major categories of CMU tasks for the first time.
Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are trained on vast amounts of text from the Internet, but do they understand the viral content that rapidly spreads online?
Approach: They introduce a dataset for CHinese Internet Meme Explanation that includes popular phrase-based memes from the Chinese Internet.
Outcome: The proposed dataset includes popular phrase-based memes from the Chinese Internet, annotated with detailed information on their meaning, origin, example sentences, types, etc.
Understanding ME? Multimodal Evaluation for Fine-grained Visual Commonsense (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models that understand image and text but also cross-reference in-between are lacking in evaluation data resources.
Approach: They propose a multimodal evaluation pipeline to automatically generate question-answer pairs to test models’ understanding of the visual scene, text, and related knowledge.
Outcome: The proposed model can answer the highly semantic VCR question correctly but fails to answer related visual question (Q2), textual question (q3), and background knowledge question ( Q4) as shallow mappings with language priors and unbalanced utilization of information between modalities.
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation approaches focus on mLLMs’ detection accuracy for binary classification tasks, which often fail to reflect the in-depth interpretive nuance of harmfulness across diverse contexts.
Approach: They propose an agent-based arena-style evaluation framework that provides context-aware and unbiased assessment for mLLMs’ understanding of multimodal harmfulness.
Outcome: The proposed framework reduces evaluation biases of judge agents and provides unbiased comparisons of mLLMs’ abilities to interpret multimodal harmfulness.
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness (2025.acl-long)

Copied to clipboard

Challenge: Existing models that assess mLLMs on harmful meme understanding are inaccurate and lack accuracy.
Approach: They propose a framework that adaptively probes the reasoning capabilities of mLLMs . their framework systematically reveals the varying performance of different target mllms a .
Outcome: The proposed framework systematically reveals the performance of different target mLLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations