Challenge: Existing datasets for humour classification are limited due to the subjectivity of the content and the multiple interpretations of the data.
Approach: They propose to annotate a multi-modal humour-annotated dataset using stand-up comedy clips and compute a humor quotient using the audience's laughter.
Outcome: The proposed scoring mechanism is validated by comparing with manual scoring methods and achieves an accuracy of 0.813 in terms of QWK.

Similar Papers

StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos (2025.findings-emnlp)

Copied to clipboard

Challenge: a new multimodal dataset of stand-up comedies is proposed to improve humor detection . the dataset is the biggest available for this type of task, and the most diverse .
Approach: They propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors.
Outcome: The proposed method improves existing models of humor detection by using audio speech recognition errors.
Telling the Whole Story: A Manually Annotated Chinese Dataset for the Analysis of Humor in Jokes (D19-1)

Copied to clipboard

Challenge: Humor plays important role in human communication, which makes it important problem for natural language processing.
Approach: They propose a novel annotation scheme to give scenarios of how humor arises in text . they report reasonable agreement between annotators and analyze the dataset .
Outcome: The proposed scheme gives scenarios of how humor arises in text . it contains key words that trigger humor, character relationship, scene, and humor categories .
Making People Laugh like a Pro: Analysing Humor Through Stand-Up Comedy (2022.lrec-1)

Copied to clipboard

Challenge: a lot of computational tools focus on standalone jokes or on occasional humorous sentences during presentations.
Approach: They propose to use stand-up comedy transcripts to extract humor from a larger narrative.
Outcome: The dataset, SCRIPTS, is built using stand-up comedy shows transcripts.
Large Dataset and Language Model Fun-Tuning for Humor Recognition (P19-1)

Copied to clipboard

Challenge: Humor recognition datasets contain only English texts and focus on puns.
Approach: They collected a dataset of jokes and funny dialogues in Russian and complemented them carefully with unfunny texts with similar lexical properties.
Outcome: The proposed method is based on the universal language model finetuning and has an F1 score of 0.91 on a test set.
When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and Its Intensity (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to generate humor using multimodal data are needed to study the role of humor in human social function.
Approach: They propose a model that automatically detects humor in the Friends TV show using multimodal data and use prerecorded laughter as annotation as it marks humor.
Outcome: The proposed model detects humor 78% of the time and how long the audience’s laughter reaction should last with a mean absolute error of 600 milliseconds.
Development and Validation of a Corpus for Machine Humor Comprehension (2020.lrec-1)

Copied to clipboard

Challenge: a Chinese humor corpus was labeled with five levels of funniness, eight skill sets of humor, and six dimensions of intent by only one annotator.
Approach: They develop a Chinese humor corpus with 3,365 jokes labeled with five levels of funniness, eight skill sets of humor, and six dimensions of intent by only one annotator.
Outcome: The proposed corpus contains 3,365 jokes from over 40 sources.
Recognizing Humour using Word Associations and Humour Anchor Extraction (C18-1)

Copied to clipboard

Challenge: Using humour anchors to improve the performance of humor recognition and interpretation is difficult for computers.
Approach: They propose to use word associations to improve humour recognition models by using humor anchors to improve the performance of semantic features.
Outcome: The proposed models improve the performance of humour recognition and interpretation tasks.
Multimodal and Multilingual Laughter Detection in Stand-Up Comedy Videos (2024.lrec-main)

Copied to clipboard

Challenge: Using TED talks, we use laughter detection software to capture humor in the sitcom genre.
Approach: They develop a multimodal multilingual dataset in Russian and English with a particular emphasis on laughter detection techniques.
Outcome: The proposed model outperforms peak detection and machine learning, while the latter shows promise and warrants further study.
The rJokes Dataset: a Large Scale Humor Collection (2020.lrec-1)

Copied to clipboard

Challenge: Humor is a complex language phenomenon that depends upon many factors, including topic, date, and recipient.
Approach: They compile a large scale humor dataset from the Reddit r/Jokes subreddit.
Outcome: The proposed dataset provides quantitative metrics for the level of humor in each joke, as determined by subreddit user feedback.
Humor Detection: A Transformer Gets the Last Laugh (D19-1)

Copied to clipboard

Challenge: Existing methods to identify humor in text have been limited to identifying humor in the text.
Approach: They propose a model that learns to identify humorous jokes based on Reddit ratings, and employ a Transformer architecture to learn from sentence context.
Outcome: The proposed model outperforms previous work on humor identification tasks with an F-measure of 93.1% for the Puns dataset and 98.6% on the Short Jokes dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations