Challenge: Large language models (LLMs) have shown strong performance in NLP tasks like text summarization and question answering.
Approach: They propose a new humor-based question-answering benchmark to assess LLMs’ reasoning through carefully crafted puns.
Outcome: Experiments on pun comprehension, resolution, and generation reveal that most LLMs struggle with generalization, even on simple tasks, consistently underperforming the human baseline.

Similar Papers

Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models for pun detection lack nuanced grasp typical of human interpretation.
Approach: They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM.
Outcome: The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns.
“A good pun is its own reword”: Can Large Language Models Understand Puns? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies on the understanding of puns in large language models (LLMs) have not explored the use of pun in creative writing and humor creation.
Approach: They propose to use pun recognition, explanation and generation tasks to evaluate the capabilities of large language models (LLMs) they adopt automated evaluation metrics from prior research and introduce new evaluation methods and metrics that align more closely with human cognition.
Outcome: The proposed methods align more closely with human cognition than previous evaluation metrics.
LLMs in Sarcasm Detection? It’s elementary! (Or is it?) (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often cited for their sophisticated pragmatic reasoning, but they collapse to random guessing on organic human speech.
Approach: They propose that LLMs have near-human competence in sarcasm detection . authors propose that this proficiency may be deceptive .
Outcome: The proposed model performance on synthetic leaderboards is a statistical mirage of competence.
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)’s influential work on the New Yorker Cartoon Caption Contest.
Approach: They propose to decompose humor understanding into three components and improve each by enhancing visual understanding through improved annotation and utilizing LLM-generated humor reasoning and explanations.
Outcome: The proposed approach achieves 82.4% accuracy in caption ranking, significantly better than the previous 67% benchmark and matches the performance of world-renowned human experts in this domain.
CANDY: Benchmarking LLMs’ Limitations and Assistive Potential in Chinese Misinformation Fact-Checking (2025.findings-emnlp)

Copied to clipboard

Challenge: CANDY is a benchmark to evaluate the capabilities and limitations of large language models (LLMs) for fact-checking misinformation.
Approach: a team of researchers develop a benchmark to evaluate the capabilities and limitations of large language models in fact-checking misinformation in Chinese.
Outcome: CANDY is a benchmark to evaluate the capabilities and limitations of large language models in fact-checking misinformation in China.
Comparing Apples to Oranges: A Dataset & Analysis of LLM Humour Understanding from Traditional Puns to Topical Jokes (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing work on humour explanation has focused on short pun-based jokes, but Large Language Models (LLMs) are not capable of generating adequate explanations of all joke types.
Approach: They compare the ability of Large Language Models (LLMs) to explain humour from simple puns to complex topical humor that requires esoteric knowledge of real-world entities and events.
Outcome: The proposed models are incapable of generating adequate explanations of all joke types, highlighting the narrow focus of most existing work on overly simple joke forms.
Blackbird language matrices (BLM), a new task for rule-like generalization in neural networks: Can Large Language Models pass the test? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to evaluate large language models for generalization lack generalization ability . current methods for evaluating LLMs are based on tests of human intelligence .
Approach: They propose to use a language task to evaluate large language models' generalisation ability . they propose to ask LLMs to solve simple variants of the RAVEN IQ test .
Outcome: The proposed task can be used to evaluate the generalisation ability of large language models . it shows that current generative models can handle the task in the sense that they understand instructions .
Challenging Large Language Models with New Tasks: A Study on their Adaptability and Robustness (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluation approaches for large language models (LLMs) rely on existing tasks and benchmarks, raising concerns about test set contamination and the genuine comprehension abilities of LLMs.
Approach: They propose to evaluate LLMs by designing new tasks, automatically generating evaluation datasets for the tasks, and conducting detailed error analyses to scrutinize LLM's adaptability to new tasks.
Outcome: The proposed method examines LLMs’ adaptability to new tasks, their sensitivity to prompt variations, and their error tendencies.
RiddleBench: A New Generative Reasoning Benchmark for LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) show remarkable capabilities, but complex reasoning skills require deeper investigation.
Approach: They propose a benchmark of 1,737 puzzles to test reasoning beyond simple pattern matching.
Outcome: The proposed model performs poorly when faced with reordered constraints or irrelevant information.
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs (2025.coling-main)

Copied to clipboard

Challenge: QUENCH is a text-based English quizzing benchmarking system for large language models (LLMs).
Approach: They propose a text-based English Quizzing Benchmark manually curated from YouTube quiz videos.
Outcome: The proposed system assesses the world knowledge and deduction capabilities of large language models via a zero-shot, open-domain quizzing setup.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations