Comparison of Pun Detection Methods Using Japanese Pun Corpus (L18-1)

Copied to clipboard

Challenge: A sampling survey of typology and component ratio analysis in Japanese puns revealed that the type of Japanese pun that had the largest proportion was a pun type with two sound sequences.
Approach: They propose a method to detect phonetically similar Japanese puns using phonological similarity and insertion / omission of prolonged sounds in addition to lexical feature.
Outcome: The proposed method is validated by adding the rule-based features to the baseline.

Similar Papers

A Survey of Pun Generation: Datasets, Evaluations and Methodologies (2025.findings-emnlp)

Copied to clipboard

Challenge: Pun generation aims to modify linguistic elements in text to produce humour or evoke double meanings.
Approach: They propose to review pun generation datasets and methods across different stages . pun generation aims to produce humour or evoke double meanings .
Outcome: This paper summarises both automated and human evaluation metrics used to assess the quality of pun generation.
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models for pun detection lack nuanced grasp typical of human interpretation.
Approach: They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM.
Outcome: The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns.
“The Boating Store Had Its Best Sail Ever”: Pronunciation-attentive Contextualized Pun Recognition (2020.acl-main)

Copied to clipboard

Challenge: Identifying and modeling puns is challenging as they involve implicit semantic or phonological tricks.
Approach: They propose a method to detect puns in a sentence and then locate them in it . they propose to capture phonetic associations between the context and phonetic symbols .
Outcome: The proposed method outperforms state-of-the-art methods in pun detection and location tasks.
Japanese Realistic Textual Entailment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 48,000 realistic examples is the largest among publicly available Japanese TE corpora . a textual entailment corpus is used to train natural language understanding . authors: to be truly helpful, machines must understand the meaning of texts.
Approach: They perform textual entailment corpus construction with 48,000 realistic examples . they use two sentences that are spontaneous or almost equivalent .
Outcome: The resulting corpus consists of 48,000 realistic Japanese examples . it is the largest among publicly available Japanese TE corpora .
Joint Detection and Location of English Puns (N19-1)

Copied to clipboard

Challenge: Existing research on puns has focused on understanding the meanings of words and phrases.
Approach: They propose a model that addresses pun detection and pun location jointly from a sequence labeling perspective.
Outcome: Empirical results show that the proposed model can handle both homographic and heterographic puns.
Building a List of Synonymous Words and Phrases of Japanese Compound Verbs (L18-1)

Copied to clipboard

Challenge: Japanese is rich in compound verbs consisting of two verbs joined together.
Approach: They built a database of Japanese "Verb + Verb" compounds semi-automatically . they extracted Japanese compound verbs from corpus and found suitable clusters .
Outcome: The proposed database extracts synonymous expressions of Japanese compound verbs from corpus . it then links the results to the "Compound Verb Lexicon"
A Unified Framework for Pun Generation with Humor Principles (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing models for generating homophonic and homographic puns lack the linguistic attributes of successful puns to resolve the split-up in existing work.
Approach: They propose a framework to generate both homophonic and homographic puns to resolve the split-up in existing works by incorporating three linguistic attributes of puns into the language models: ambiguity, distinctiveness, and surprise.
Outcome: The proposed model over strong baselines shows that it can generate both homophonic and homographic puns.
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
Developing Japanese CLIP Models Leveraging an Open-weight LLM for Large-scale Dataset Translation (2025.naacl-srw)

Copied to clipboard

Challenge: lack of large-scale open Japanese image-text pairs poses a significant barrier to the development of vision-language models.
Approach: They construct large-scale Japanese image-text pairs using machine translation and pre-trained CLIP models on a Japanese dataset.
Outcome: The results show that pre-trained models achieve competitive average scores on Japanese culture tasks compared to models of similar size.
A unified approach to sentence segmentation of punctuated text in many languages (2021.acl-long)

Copied to clipboard

Challenge: Existing tools for segmenting punctuated text in many languages are limited in their language coverage and evaluation is ad hoc.
Approach: They propose a new context-based modeling approach that can be trained on noisily-annotated data.
Outcome: The proposed model exceeds baselines set by existing methods on English corpora and performs well on average on new multilingual evaluation set.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations