Challenge: Existing Vision-Language models perform poorly on satirical image detecting tasks . satire and humor are powerful tools to highlight issues, provoke thought, and encourage critical perspective .
Approach: They propose to use a dataset to evaluate satirical images and satire images to detect satiric images . they also propose to generate the reason behind the image being satiral by generating one half of the image to be satisfying .
Outcome: The proposed dataset contains 2547 images, 1084 satirical and 1463 non-satirically, with different artistic styles.

Similar Papers

Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)

Copied to clipboard

Challenge: Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny.
Approach: They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny .
Outcome: The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny.
Caption Enriched Samples for Improving Hateful Memes Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy.
Approach: They propose to use an off-the-shelf caption generator to capture the first image and overlayed text.
Outcome: The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text .
ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos (2023.emnlp-main)

Copied to clipboard

Challenge: despite its importance, there are few datasets that cover multimodal counterfactual reasoning . a dataset focusing on this area is limited because of its limited coverage over synthetic environments .
Approach: They develop a video question answering dataset that provides questions on multimodal reasoning . they ask questions about counterfactual hypotheses over visual events .
Outcome: The proposed dataset shows a significant performance gap between models and humans . it provides questions that span physical, social, and temporal dimensions .
Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue.
Approach: They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments.
Outcome: The proposed dataset improves the ability of large language models to understand humor.
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos (2025.findings-emnlp)

Copied to clipboard

Challenge: a new multimodal dataset of stand-up comedies is proposed to improve humor detection . the dataset is the biggest available for this type of task, and the most diverse .
Approach: They propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors.
Outcome: The proposed method improves existing models of humor detection by using audio speech recognition errors.
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor (D19-1)

Copied to clipboard

Challenge: Humor is a unique and creative communicative behavior often displayed during social interactions.
Approach: They present a dataset that allows to model multimodal language used in expressing humor using text, visual and acoustic communication.
Outcome: The proposed framework opens the door to understanding multimodal language used in expressing humor.
Punny Captions: Witty Wordplay in Image Descriptions (N18-2)

Copied to clipboard

Challenge: Developing computational models that can produce contextually witty image descriptions is challenging because of the large corpus of sentences that are not available for large scale corpora.
Approach: They propose to use linguistic wordplay, specifically puns, to generate witty image descriptions from large corpus of sentences or encode them via an encoder-decoder neural network architecture.
Outcome: The proposed models perform better than baseline models using human data and show that they are slightly wittier than human-written witty descriptions.
SarcNet: A Multilingual Multimodal Sarcasm Detection Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Sarcasm is an implicit form of sarcasm, involving an intended meaning that contradicts the literal expression . human use conflict between factual information and a statement as cues to detect sarcasm . sarkasmatic analysis is challenging due to its implicit nature .
Approach: They propose a multimodal sarcasm detection dataset that uses multiple modalities to detect sarcasm.
Outcome: The proposed model improves on previous models based on a single label . human sarcasm cannot be detected using a unified label across multiple modalities .
What A Sunny Day ☔: Toward Emoji-Sensitive Irony Detection (D19-55)

Copied to clipboard

Challenge: Existing datasets for irony detection only contain 10% of ironic tweets with emojis . 45% of internet users in the united states use an e-moji in social media .
Approach: They propose to use emojis to analyze irony detection datasets to train classifiers.
Outcome: The proposed pipeline can be used to analyze irony detection datasets using emojis.
FanChuan: A Multilingual and Graph-Structured Benchmark For Parody Detection and Analysis (2025.findings-acl)

Copied to clipboard

Challenge: Parody is an emerging phenomenon on social media, where individuals imitate a role or position opposite to their own . limited available data and deficient diversity in current datasets hinder study of parody .
Approach: They build a dataset of parody users and annotated comments from both English and Chinese corpora to test parody detection and comment sentiment analysis.
Outcome: The proposed datasets provide richer contextual information, which is lacking in existing datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations