Papers with humor
Transfer Learning for Humor Detection by Twin Masked Yellow Muppets (2022.aacl-short)
Copied to clipboard
| Challenge: | Existing humor classification systems have been dealing with different forms of humor independently. |
| Approach: | They propose to combine different forms of humor to tackle different humor types by a shared-private multitask architecture using a transfer learning paradigm. |
| Outcome: | The proposed architecture shows statistically significant improvements over baselines and accounting for new state-of-the-art figures for two datasets. |
Understanding Figurative Meaning through Explainable Visual Entailment (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing models for visual entailment and visual question-answering have limited ability to understand figurative meaning in images and captions. |
| Approach: | They propose a task framing the figurative meaning understanding problem as an explainable visual entailment task where the model has to predict whether the image entitles a caption and justify the predicted label with a textual explanation. |
| Outcome: | The proposed dataset contains 6,027 image, caption, label, explanation instances covering five diverse figurative phenomena. |
Towards Generation and Recognition of Humorous Texts in Portuguese (2023.eacl-srw)
Copied to clipboard
| Challenge: | This PhD thesis focuses on the automatic generation and recognition of verbal punning humor in Portuguese. |
| Approach: | They propose to combine natural language generation and cognitive processing to generate and recognize verbal humor in Portuguese. |
| Outcome: | The proposed methods aim to generate and recognize humor in Portuguese, an underdeveloped language compared to English. |
Humor Recognition Using Deep Learning (N18-2)
Copied to clipboard
| Challenge: | Humor is an essential but most fascinating element in personal communication. |
| Approach: | They propose a convolutional neural network with extensive filter size and filter number to increase the depth of networks. |
| Outcome: | The proposed model outperforms existing models on accuracy, precision and recall . the proposed model can learn to distinguish between humorous and nonhumorous texts . |
Crossing the Line: Where do Demographic Variables Fit into Humor Detection? (2020.acl-srw)
Copied to clipboard
| Challenge: | Recent shared tasks for humor classification have struggled with two issues: the data comprises a highly constrained genre of humor which does not broadly represent humor, or the data is so indiscriminate that the inter-annotator agreement on its humor content is drastically low. |
| Approach: | They propose adding demographic information about the humor annotators in order to bin ratings more sensibly and adding an ‘offensive’ label to distinguish between different generations, in terms of humor. |
| Outcome: | The proposed system could be adapted to more nuanced tasks and improve performance on downstream tasks, such as content moderation. |
Stimulating Creativity with FunLines: A Case Study of Humor Generation in Headlines (2020.acl-demos)
Copied to clipboard
| Challenge: | FunLines is an online game that allows players to generate and rate funny news headlines . it is difficult to generate data that depends on human creativity, and measuring creativity often requires more effort. |
| Approach: | They propose a game where players edit news headlines to make them funny and rate the funniness of headlines edited by others. |
| Outcome: | The proposed game outperforms other crowdsourcing approaches in generating humor datasets. |
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs (2026.eacl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide strong generative capabilities, but many applications require explicit and fine-grained control over specific textual concepts. |
| Approach: | They propose a framework for fine-grained controllability for single- and dual-concept scenarios . they find performance drops in the dual-constituency setting, even though chosen concepts should be separable . |
| Outcome: | The proposed framework shows that models struggle with compositionality even when concepts are intuitively independent. |
“The Boating Store Had Its Best Sail Ever”: Pronunciation-attentive Contextualized Pun Recognition (2020.acl-main)
Copied to clipboard
| Challenge: | Identifying and modeling puns is challenging as they involve implicit semantic or phonological tricks. |
| Approach: | They propose a method to detect puns in a sentence and then locate them in it . they propose to capture phonetic associations between the context and phonetic symbols . |
| Outcome: | The proposed method outperforms state-of-the-art methods in pun detection and location tasks. |
Combining Humor and Sarcasm for Improving Political Parody Detection (2022.naacl-main)
Copied to clipboard
| Challenge: | Parody is a figurative device used for mimicking entities for comedic or critical purposes. |
| Approach: | They propose a multi-encoder model that combines three parallel encoders to enrich parody-specific representations with humor and sarcasm information. |
| Outcome: | The proposed model outperforms state-of-the-art methods on a dataset of political parody tweets. |
Mining Effective Features Using Quantum Entropy for Humor Recognition (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies on humor recognition do not understand the mechanisms that generate humor. |
| Approach: | They propose to use quantum entropy to represent the semantic uncertainty of the setup and punchline as features for humor recognition. |
| Outcome: | The proposed features are more effective than baselines for recognizing humorous and non-humorous texts on the SemEval2021 task 7 dataset. |
Exploiting Syntactic Structures for Humor Recognition (C18-1)
Copied to clipboard
| Challenge: | Using syntactic structure features, we find humor recognition is a kind of style . |
| Approach: | They propose to exploit syntactic structure features to enhance humor recognition . they find syntastic structure features consistently correlate with humor . |
| Outcome: | The proposed method achieves significant improvements compared with baselines. |
Can Language Models Laugh at YouTube Short-form Videos? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that focus on verbal cues and focus on short-form funny videos focus on focusing on verbs and visual cue. |
| Approach: | They curate a user-generated dataset of 10K multimodal funny videos from YouTube and annotate each video with timestamps and explanations for funny moments. |
| Outcome: | The proposed dataset improves the ability of large language models to understand humor. |
ExPUNations: Augmenting Puns with Keywords and Explanations (2022.emnlp-main)
Copied to clipboard
Jiao Sun, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Tagyoung Chung, Jing Huang, Yang Liu, Nanyun Peng
| Challenge: | Puns add the challenge of fusing commonsense and world knowledge with the ability to interpret lexical-semantic ambiguity. |
| Approach: | They propose to augment existing datasets with detailed crowdsourced annotations of puns, keywords and fine-grained funniness ratings to challenge current models' ability to understand and generate humor. |
| Outcome: | The proposed tasks include explanation generation to aid with pun classification and keyword-conditioned pun generation to challenge state-of-the-art models' ability to understand and generate humor. |
“I See What You Did There”: Can Large Vision-Language Models Understand Multimodal Puns? (2026.acl-long)
Copied to clipboard
Naen Xu, Jiayi Sheng, Changjiang Li, Chunyi Zhou, Yuyuan Li, Tianyu Du, Jun Wang, Zhihui Fu, Jinbao Li, Shouling Ji
| Challenge: | Puns are a common form of rhetorical wordplay that exploits polysemy and phonetic similarity to create humor. |
| Approach: | They propose a multimodal pun generation pipeline and a model to evaluate their understanding of puns. |
| Outcome: | The proposed benchmark improves the understanding of multimodal puns by 16.5% in the F1 test. |
Making People Laugh like a Pro: Analysing Humor Through Stand-Up Comedy (2022.lrec-1)
Copied to clipboard
| Challenge: | a lot of computational tools focus on standalone jokes or on occasional humorous sentences during presentations. |
| Approach: | They propose to use stand-up comedy transcripts to extract humor from a larger narrative. |
| Outcome: | The dataset, SCRIPTS, is built using stand-up comedy shows transcripts. |
A Sentiment and Emotion Aware Multimodal Multiparty Humor Recognition in Multilingual Conversational Setting (2022.coling-1)
Copied to clipboard
| Challenge: | Humor is an essential aspect of daily conversation, and people try to provoke humor in their talks. |
| Approach: | They propose a multitask framework that annotates Hindi utterances with sentiment and emotion classes. |
| Outcome: | The proposed framework improves on the recently released Hindi Humor dataset . it takes sentiment and emotion into account to understand humor . |
Telling the Whole Story: A Manually Annotated Chinese Dataset for the Analysis of Humor in Jokes (D19-1)
Copied to clipboard
| Challenge: | Humor plays important role in human communication, which makes it important problem for natural language processing. |
| Approach: | They propose a novel annotation scheme to give scenarios of how humor arises in text . they report reasonable agreement between annotators and analyze the dataset . |
| Outcome: | The proposed scheme gives scenarios of how humor arises in text . it contains key words that trigger humor, character relationship, scene, and humor categories . |
The rJokes Dataset: a Large Scale Humor Collection (2020.lrec-1)
Copied to clipboard
| Challenge: | Humor is a complex language phenomenon that depends upon many factors, including topic, date, and recipient. |
| Approach: | They compile a large scale humor dataset from the Reddit r/Jokes subreddit. |
| Outcome: | The proposed dataset provides quantitative metrics for the level of humor in each joke, as determined by subreddit user feedback. |
FanChuan: A Multilingual and Graph-Structured Benchmark For Parody Detection and Analysis (2025.findings-acl)
Copied to clipboard
Yilun Zheng, Sha Li, Fangkun Wu, Yang Ziyi, Lin Hongchao, Zhichao Hu, Cai Xinjun, Ziming Wang, Jinxuan Chen, Sitao Luan, Jiahao Xu, Lihui Chen
| Challenge: | Parody is an emerging phenomenon on social media, where individuals imitate a role or position opposite to their own . limited available data and deficient diversity in current datasets hinder study of parody . |
| Approach: | They build a dataset of parody users and annotated comments from both English and Chinese corpora to test parody detection and comment sentiment analysis. |
| Outcome: | The proposed datasets provide richer contextual information, which is lacking in existing datasets. |
So Hateful! Building a Multi-Label Hate Speech Annotated Arabic Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Social media enables widespread propagation of hate speech targeting groups based on ethnicity, religion, or other characteristics. |
| Approach: | They analyze 70,000 Arabic tweets to identify hate speech patterns and train models . 15% of tweets contain offensive language while 6% have hate speech . authors hope to prevent spread of hateful content on social media platforms . |
| Outcome: | The analysis of 70,000 Arabic tweets shows that 15% of tweets contain offensive language while 6% have hate speech . 10% of tweet provide verifiable factual claims, and 7% are deemed important . |
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound (2026.acl-long)
Copied to clipboard
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Song-Chun Zhu, Bo Zhao, Zilong Zheng
| Challenge: | Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos. |
| Approach: | They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues. |
| Outcome: | The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor . |
Investigating Counterfactual Unfairness in LLMs towards Identities through Humor (2026.acl-long)
Copied to clipboard
Shubin Kim, Yejin Son, Junyeong Park, Keummin Ka, Seungbeen Lee, Jaeyoung Lee, Hyeju Jang, Alice Oh, Youngjae Yu
| Challenge: | Large Language Models (LLMs) absorb social and cultural biases embedded in vast web-scale corpora and are increasingly deployed in high-stakes domains such as hiring, education, and law. |
| Approach: | They propose a framework to investigate counterfactual unfairness through humor by observing how the model’s responses change when we swap who speaks and who is addressed while holding other factors constant. |
| Outcome: | The proposed framework covers humor generation refusal, speaker intention inference, and relational/societal impact prediction tasks. |