| Challenge: | A sampling survey of typology and component ratio analysis in Japanese puns revealed that the type of Japanese pun that had the largest proportion was a pun type with two sound sequences. |
| Approach: | They propose a method to detect phonetically similar Japanese puns using phonological similarity and insertion / omission of prolonged sounds in addition to lexical feature. |
| Outcome: | The proposed method is validated by adding the rule-based features to the baseline. |
Similar Papers
A Survey of Pun Generation: Datasets, Evaluations and Methodologies (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Pun generation aims to modify linguistic elements in text to produce humour or evoke double meanings. |
| Approach: | They propose to review pun generation datasets and methods across different stages . pun generation aims to produce humour or evoke double meanings . |
| Outcome: | This paper summarises both automated and human evaluation metrics used to assess the quality of pun generation. |
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)
Copied to clipboard
Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar, Jose Camacho-Collados
| Challenge: | Existing models for pun detection lack nuanced grasp typical of human interpretation. |
| Approach: | They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM. |
| Outcome: | The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns. |
“The Boating Store Had Its Best Sail Ever”: Pronunciation-attentive Contextualized Pun Recognition (2020.acl-main)
Copied to clipboard
| Challenge: | Identifying and modeling puns is challenging as they involve implicit semantic or phonological tricks. |
| Approach: | They propose a method to detect puns in a sentence and then locate them in it . they propose to capture phonetic associations between the context and phonetic symbols . |
| Outcome: | The proposed method outperforms state-of-the-art methods in pun detection and location tasks. |
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts. |
| Approach: | They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent . |
| Outcome: | The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora . |
Joint Detection and Location of English Puns (N19-1)
Copied to clipboard
| Challenge: | Existing research on puns has focused on understanding the meanings of words and phrases. |
| Approach: | They propose a model that addresses pun detection and pun location jointly from a sequence labeling perspective. |
| Outcome: | Empirical results show that the proposed model can handle both homographic and heterographic puns. |
Building a List of Synonymous Words and Phrases of Japanese Compound Verbs (L18-1)
Copied to clipboard
| Challenge: | Japanese is rich in compound verbs consisting of two verbs joined together. |
| Approach: | They built a database of Japanese "Verb + Verb" compounds semi-automatically . they extracted Japanese compound verbs from corpus and found suitable clusters . |
| Outcome: | The proposed database extracts synonymous expressions of Japanese compound verbs from corpus . it then links the results to the "Compound Verb Lexicon" |
A Unified Framework for Pun Generation with Humor Principles (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for generating homophonic and homographic puns lack the linguistic attributes of successful puns to resolve the split-up in existing work. |
| Approach: | They propose a framework to generate both homophonic and homographic puns to resolve the split-up in existing works by incorporating three linguistic attributes of puns into the language models: ambiguity, distinctiveness, and surprise. |
| Outcome: | The proposed model over strong baselines shows that it can generate both homophonic and homographic puns. |
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)
Copied to clipboard
Hanae Koiso, Yasuharu Den, Yuriko Iseki, Wakako Kashino, Yoshiko Kawabata, Ken’ya Nishikawa, Yayoi Tanaka, Yasuyuki Usuda
| Challenge: | a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations . |
| Approach: | They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner. |
| Outcome: | The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings. |
Developing Japanese CLIP Models Leveraging an Open-weight LLM for Large-scale Dataset Translation (2025.naacl-srw)
Copied to clipboard
| Challenge: | lack of large-scale open Japanese image-text pairs poses a significant barrier to the development of vision-language models. |
| Approach: | They construct large-scale Japanese image-text pairs using machine translation and pre-trained CLIP models on a Japanese dataset. |
| Outcome: | The results show that pre-trained models achieve competitive average scores on Japanese culture tasks compared to models of similar size. |
A unified approach to sentence segmentation of punctuated text in many languages (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools for segmenting punctuated text in many languages are limited in their language coverage and evaluation is ad hoc. |
| Approach: | They propose a new context-based modeling approach that can be trained on noisily-annotated data. |
| Outcome: | The proposed model exceeds baselines set by existing methods on English corpora and performs well on average on new multilingual evaluation set. |