Papers with Phunny
“What do you call a dog that is incontrovertibly true? Dogma”: Testing LLM Generalization through Humor (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown strong performance in NLP tasks like text summarization and question answering. |
| Approach: | They propose a new humor-based question-answering benchmark to assess LLMs’ reasoning through carefully crafted puns. |
| Outcome: | Experiments on pun comprehension, resolution, and generation reveal that most LLMs struggle with generalization, even on simple tasks, consistently underperforming the human baseline. |