Challenge: In this paper, we analyze memes as a form of language subject to the same kinds of sociolinguistic variation as other modalities, such as written language and speech.
Approach: They propose a computational pipeline to cluster memes into templates and semantic variables, taking advantage of their multimodal structure to learn meme semantics from an unstructured dataset.
Outcome: The proposed method uses 3.8M images from a reddit meme database to analyze linguistic variation in memes.

Similar Papers

Characterizing English Variation across Social Media Communities with BERT (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies characterizing language variation across Internet social groups have focused on the types of words used by these groups.
Approach: They extend this study by employing BERT to characterize variation in the senses of words as well, analyzing two months of English comments in 474 Reddit communities.
Outcome: The proposed study analyzes two months of English comments in 474 Reddit communities and ties language variation with community behavior.
Computational Meme Understanding: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Computational Meme Understanding (CMU) is a collection of tasks involving the automated comprehension of memes.
Approach: They propose a comprehensive taxonomy for memes along three dimensions – forms, functions, and topics and introduce three key tasks for Computational Meme Understanding, namely classification, interpretation, and explanation.
Outcome: The proposed model is based on a taxonomy of memes along three dimensions and is compared to existing models and datasets.
MemoSen: A Multimodal Dataset for Sentiment Analysis of Memes (2022.lrec-1)

Copied to clipboard

Challenge: Recent studies on sentiment analysis of memes have focused on English, but there is a significant barrier to performing multimodal sentiment analysis research in resource-constrained languages like Bengali.
Approach: They propose to use a Bengali dataset to perform multimodal sentiment analysis in low resource languages.
Outcome: The proposed dataset for Bengali contains 4417 memes with three annotated labels positive, negative, and neutral.
MetaMeme: A Dataset for Meme Template and Meta-Category Classification (2025.naacl-srw)

Copied to clipboard

Challenge: a new dataset for classifying memes by their template and communicative intent is presented.
Approach: They propose a new dataset for classifying memes by their template and communicative intent.
Outcome: The proposed method outperforms existing methods in classifying memes by their template and communicative intent.
FigMemes: A Dataset for Figurative Language Identification in Politically-Opinionated Memes (2022.emnlp-main)

Copied to clipboard

Challenge: FigMemes is a dataset for figurative language classification in politically-opinionated memes.
Approach: They propose to use figurative language classification to identify politically-opinionated memes by analyzing their datasets and comparing them to other machine learning models.
Outcome: The proposed dataset includes annotations of six commonly used types of figurative language in politically-opinionated memes and a wide range of topics and visual styles.
Variationist: Exploring Multifaceted Variation and Bias in Written Language Data (2024.acl-demos)

Copied to clipboard

Challenge: Existing tools that inspect and visualize language data are limited in their capabilities.
Approach: They propose a highly-modular, extensible, and task-agnostic tool that inspects language variation and bias across multiple variables, language units, and diverse metrics.
Outcome: The proposed tool can inspect and visualize language variation and bias across variables, language units, and diverse metrics that go beyond descriptive statistics.
MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization (2023.acl-long)

Copied to clipboard

Challenge: Besides digital archiving of memes and their metadata, there is no efficient way to deduce a meme’s context dynamically.
Approach: They propose a task to mine the context that succinctly explains the background of a meme and a related document to capture cross-modal semantic dependencies between the meme and the context.
Outcome: The proposed dataset outperforms existing systems and shows that it can capture cross-modal semantic dependencies between the meme and the context.
Meme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations (2024.eacl-long)

Copied to clipboard

Challenge: Recent laws like “right to explanations” have spurred research in developing interpretable models . a recent study has shown that multimodal explanations improve performance in generating textual justifications .
Approach: They propose to use visual and textual modalities to explain why a given meme is cyberbullying . they use a Contrastive Language-Image Pretraining approach to generate textual justifications .
Outcome: The proposed model improves performance in visual and textual explanations and identifies the visual evidence supporting a decision.
MemeCap: A Dataset for Captioning and Interpreting Memes (2023.emnlp-main)

Copied to clipboard

Challenge: a new dataset aims to understand meme captioning tasks using visual metaphors . vision and language models are proving to be effective in image captioning and visual question answering tasks .
Approach: They present a dataset that contains 6.3K memes and 6.3k meme captions . they show that vision and language models still struggle with visual metaphors despite their advanced capabilities .
Outcome: The proposed dataset contains 6.3K memes along with the title of the post containing the meme, meme captions, literal image caption, and visual metaphors.
Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are trained on vast amounts of text from the Internet, but do they understand the viral content that rapidly spreads online?
Approach: They introduce a dataset for CHinese Internet Meme Explanation that includes popular phrase-based memes from the Chinese Internet.
Outcome: The proposed dataset includes popular phrase-based memes from the Chinese Internet, annotated with detailed information on their meaning, origin, example sentences, types, etc.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations