Challenge: a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text .
Approach: They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags.
Outcome: The proposed method detects humor in code-mixed tweets in English-Hindi.

Similar Papers

UR-FUNNY: A Multimodal Language Dataset for Understanding Humor (D19-1)

Copied to clipboard

Challenge: Humor is a unique and creative communicative behavior often displayed during social interactions.
Approach: They present a dataset that allows to model multimodal language used in expressing humor using text, visual and acoustic communication.
Outcome: The proposed framework opens the door to understanding multimodal language used in expressing humor.
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)

Copied to clipboard

Challenge: a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people.
Approach: They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook.
Outcome: The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.
Towards Automated Semantic Role Labelling of Hindi-English Code-Mixed Tweets (D19-55)

Copied to clipboard

Challenge: a new system for semantic role labelling of Hindi-English code-mixed tweets is proposed . code-mixing is a largely observed phenomenon in colloquial usage and on social media .
Approach: They propose a system for automating Semantic Role Labelling of Hindi-English code-mixed tweets.
Outcome: The proposed system gives an overall accuracy of 84% for Argument Classification, a 10% increase over the existing rule-based model.
Corpus Creation and Emotion Prediction for Hindi-English Code-Mixed Social Media Text (N18-4)

Copied to clipboard

Challenge: Emotion Prediction is a natural language processing task dealing with detection and classification of emotions in monolingual and bilingual texts.
Approach: They propose a machine learning system which uses various machine learning techniques to detect emotion associated with tweets.
Outcome: The proposed system uses various machine learning techniques to detect emotion associated with the text.
Revealing the impact of synthetic native samples and multi-tasking strategies in Hindi-English code-mixed humour and sarcasm detection (2025.findings-emnlp)

Copied to clipboard

Challenge: Specifically, we tried native sample mixing, multi-task learning, and prompting and instruction finetuning very large multilingual language models (VMLMs).
Approach: They used native sample mixing, multi-task learning and prompting and instruction finetuning to improve code-mixed humour and sarcasm detection.
Outcome: The proposed methods improve humour and sarcasm detection by adding native samples to training sets and multitask learning and prompting and instruction finetuning VMLMs.
Corpora and Baselines for Humour Recognition in Portuguese (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on the recognition of verbal humour in Portuguese has not been done . humor recognition is a sign of fluency in a language, and is not yet widely used in other languages.
Approach: They propose to create three corpora covering two styles of humour and four sources of non-humorous text that are used for testing computational models.
Outcome: The proposed models can be used to train and test models in Portuguese, and may be used as baselines for future projects.
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos (2025.findings-emnlp)

Copied to clipboard

Challenge: a new multimodal dataset of stand-up comedies is proposed to improve humor detection . the dataset is the biggest available for this type of task, and the most diverse .
Approach: They propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors.
Outcome: The proposed method improves existing models of humor detection by using audio speech recognition errors.
Large Dataset and Language Model Fun-Tuning for Humor Recognition (P19-1)

Copied to clipboard

Challenge: Humor recognition datasets contain only English texts and focus on puns.
Approach: They collected a dataset of jokes and funny dialogues in Russian and complemented them carefully with unfunny texts with similar lexical properties.
Outcome: The proposed method is based on the universal language model finetuning and has an F1 score of 0.91 on a test set.
HAHA 2019 Dataset: A Corpus for Humor Analysis in Spanish (2020.lrec-1)

Copied to clipboard

Challenge: 30,000 Spanish tweets were crowd-annotated with humor value and funniness score . the corpus contains approximately 38.6% of humorous tweets with an average score of 2.04 in a scale from 1 to 5 for the humorous tweet.
Approach: They develop a corpus of 30,000 Spanish tweets crowd-annotated with humor value and funniness score.
Outcome: The results obtained from the 30,000 tweets in the Spanish language are encouraging.
Leveraging Social Context for Humor Recognition and Sense of Humor Evaluation in Social Media with a New Chinese Humor Corpus - HumorWB (2024.lrec-main)

Copied to clipboard

Challenge: Existing humor computing research focuses on content while neglecting interaction relationships in social media.
Approach: They propose a dataset which introduces social context information from social media . they propose 'humor recognition' task and 'horror evaluation task'
Outcome: The proposed model incorporates social context information from social media . it shows that it is efficient and can be used to evaluate humor in real life .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations