Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work on humour explanation has focused on short pun-based jokes, but Large Language Models (LLMs) are not capable of generating adequate explanations of all joke types. |
| Approach: | They compare the ability of Large Language Models (LLMs) to explain humour from simple puns to complex topical humor that requires esoteric knowledge of real-world entities and events. |
| Outcome: | The proposed models are incapable of generating adequate explanations of all joke types, highlighting the narrow focus of most existing work on overly simple joke forms. |
Similar Papers
“A good pun is its own reword”: Can Large Language Models Understand Puns? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on the understanding of puns in large language models (LLMs) have not explored the use of pun in creative writing and humor creation. |
| Approach: | They propose to use pun recognition, explanation and generation tasks to evaluate the capabilities of large language models (LLMs) they adopt automated evaluation metrics from prior research and introduce new evaluation methods and metrics that align more closely with human cognition. |
| Outcome: | The proposed methods align more closely with human cognition than previous evaluation metrics. |
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)
Copied to clipboard
Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar, Jose Camacho-Collados
| Challenge: | Existing models for pun detection lack nuanced grasp typical of human interpretation. |
| Approach: | They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM. |
| Outcome: | The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns. |
Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme Explanation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are trained on vast amounts of text from the Internet, but do they understand the viral content that rapidly spreads online? |
| Approach: | They introduce a dataset for CHinese Internet Meme Explanation that includes popular phrase-based memes from the Chinese Internet. |
| Outcome: | The proposed dataset includes popular phrase-based memes from the Chinese Internet, annotated with detailed information on their meaning, origin, example sentences, types, etc. |
Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue. |
| Approach: | They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments. |
| Outcome: | The proposed dataset improves the ability of large language models to understand humor. |
The rJokes Dataset: a Large Scale Humor Collection (2020.lrec-1)
Copied to clipboard
| Challenge: | Humor is a complex language phenomenon that depends upon many factors, including topic, date, and recipient. |
| Approach: | They compile a large scale humor dataset from the Reddit r/Jokes subreddit. |
| Outcome: | The proposed dataset provides quantitative metrics for the level of humor in each joke, as determined by subreddit user feedback. |
Recognizing Humour using Word Associations and Humour Anchor Extraction (C18-1)
Copied to clipboard
| Challenge: | Using humour anchors to improve the performance of humor recognition and interpretation is difficult for computers. |
| Approach: | They propose to use word associations to improve humour recognition models by using humor anchors to improve the performance of semantic features. |
| Outcome: | The proposed models improve the performance of humour recognition and interpretation tasks. |
When the LM misunderstood the human chuckled: Analyzing garden path effects in humans and language models (2025.acl-long)
Copied to clipboard
| Challenge: | Modern Large Language Models (LLMs) have shown human-like abilities in many language tasks, sparking interest in comparing LLMs’ and humans’ language processing. |
| Approach: | They propose to answer two questions: 1. What makes garden-path sentences hard for humans? 2. Do the same reasons make garden- path sentences hard? |
| Outcome: | The proposed models show that humans struggle with specific syntactic complexities, with some models showing high correlation with human comprehension. |
“What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through Humor (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown strong performance in NLP tasks like text summarization and question answering. |
| Approach: | They propose a new humor-based question-answering benchmark to assess LLMs’ reasoning through carefully crafted puns. |
| Outcome: | Experiments on pun comprehension, resolution, and generation reveal that most LLMs struggle with generalization, even on simple tasks, consistently underperforming the human baseline. |
Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)
Copied to clipboard
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, Yejin Choi
| Challenge: | Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny. |
| Approach: | They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny . |
| Outcome: | The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny. |
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Slang is a commonly used type of informal language that poses a daunting challenge to NLP systems. |
| Approach: | They compare human-attested slang and swiss-generated slurs with machine-generated ones . they find that LLMs have significant knowledge about the creative aspects of sling . |
| Outcome: | The proposed model compares human and machine-generated slang usages to find biases in human perceptions of sling . the results suggest that human-attested slms have significant knowledge about the creative aspects of a language . |