Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)
Copied to clipboard
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, Yejin Choi
| Challenge: | Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny. |
| Approach: | They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny . |
| Outcome: | The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny. |
Similar Papers
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics (2025.findings-emnlp)
Copied to clipboard
| Challenge: | PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences. |
| Approach: | PixelHumor is a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs’ ability to interpret multimodal humor and recognize narrative sequences. |
| Outcome: | Experiments with state-of-the-art LMMs reveal that top models achieve only 61% accuracy in panel sequencing, far below human performance. |
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs (2025.findings-emnlp)
Copied to clipboard
Kuan Lok Zhou, Jiayi Chen, Siddharth Suresh, Reuben Narad, Timothy T. Rogers, Lalit K Jain, Robert D Nowak, Bob Mankoff, Jifan Zhang
| Challenge: | Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)’s influential work on the New Yorker Cartoon Caption Contest. |
| Approach: | They propose to decompose humor understanding into three components and improve each by enhancing visual understanding through improved annotation and utilizing LLM-generated humor reasoning and explanations. |
| Outcome: | The proposed approach achieves 82.4% accuracy in caption ranking, significantly better than the previous 67% benchmark and matches the performance of world-renowned human experts in this domain. |
Recognizing Humour using Word Associations and Humour Anchor Extraction (C18-1)
Copied to clipboard
| Challenge: | Using humour anchors to improve the performance of humor recognition and interpretation is difficult for computers. |
| Approach: | They propose to use word associations to improve humour recognition models by using humor anchors to improve the performance of semantic features. |
| Outcome: | The proposed models improve the performance of humour recognition and interpretation tasks. |
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on humour explanation has focused on short pun-based jokes, but Large Language Models (LLMs) are not capable of generating adequate explanations of all joke types. |
| Approach: | They compare the ability of Large Language Models (LLMs) to explain humour from simple puns to complex topical humor that requires esoteric knowledge of real-world entities and events. |
| Outcome: | The proposed models are incapable of generating adequate explanations of all joke types, highlighting the narrow focus of most existing work on overly simple joke forms. |
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound (2026.acl-long)
Copied to clipboard
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Song-Chun Zhu, Bo Zhao, Zilong Zheng
| Challenge: | Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos. |
| Approach: | They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues. |
| Outcome: | The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor . |
Computational Meme Understanding: A Survey (2024.emnlp-main)
Copied to clipboard
| Challenge: | Computational Meme Understanding (CMU) is a collection of tasks involving the automated comprehension of memes. |
| Approach: | They propose a comprehensive taxonomy for memes along three dimensions – forms, functions, and topics and introduce three key tasks for Computational Meme Understanding, namely classification, interpretation, and explanation. |
| Outcome: | The proposed model is based on a taxonomy of memes along three dimensions and is compared to existing models and datasets. |
SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Despite advances in artificial intelligence, building social intelligence remains a challenge. |
| Approach: | They propose a task to explain why people laugh in a video and a dataset to do this. |
| Outcome: | The proposed dataset generates plausible explanations for laughter in video and in-the-wild videos. |
Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue. |
| Approach: | They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments. |
| Outcome: | The proposed dataset improves the ability of large language models to understand humor. |
Humor Detection: A Transformer Gets the Last Laugh (D19-1)
Copied to clipboard
| Challenge: | Existing methods to identify humor in text have been limited to identifying humor in the text. |
| Approach: | They propose a model that learns to identify humorous jokes based on Reddit ratings, and employ a Transformer architecture to learn from sentence context. |
| Outcome: | The proposed model outperforms previous work on humor identification tasks with an F-measure of 93.1% for the Puns dataset and 98.6% on the Short Jokes dataset. |
Caption Enriched Samples for Improving Hateful Memes Detection (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy. |
| Approach: | They propose to use an off-the-shelf caption generator to capture the first image and overlayed text. |
| Outcome: | The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text . |