Puntuguese: A Corpus of Puns in Portuguese with Micro-edits (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpus of punning humor in Portuguese is unfit for machine learning due to data leakage.
Approach: They propose to use Puntuguese to create a corpus of punning humor in Portuguese that is significantly more difficult to recognize than the previous corpus.
Outcome: The proposed corpus achieves an F1-Score of 68.9% and is significantly more difficult than the previous corpus.

Similar Papers

Towards Generation and Recognition of Humorous Texts in Portuguese (2023.eacl-srw)

Copied to clipboard

Challenge: This PhD thesis focuses on the automatic generation and recognition of verbal punning humor in Portuguese.
Approach: They propose to combine natural language generation and cognitive processing to generate and recognize verbal humor in Portuguese.
Outcome: The proposed methods aim to generate and recognize humor in Portuguese, an underdeveloped language compared to English.
Corpora and Baselines for Humour Recognition in Portuguese (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on the recognition of verbal humour in Portuguese has not been done . humor recognition is a sign of fluency in a language, and is not yet widely used in other languages.
Approach: They propose to create three corpora covering two styles of humour and four sources of non-humorous text that are used for testing computational models.
Outcome: The proposed models can be used to train and test models in Portuguese, and may be used as baselines for future projects.
Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing (2020.coling-main)

Copied to clipboard

Challenge: Existing corpus for automatic post-editing of English and Brazilian Portuguese is limited.
Approach: They introduce a corpus for Automatic Post-Editing of English and Brazilian Portuguese.
Outcome: The proposed corpus improves on the English and Brazilian Portuguese languages.
Large Dataset and Language Model Fun-Tuning for Humor Recognition (P19-1)

Copied to clipboard

Challenge: Humor recognition datasets contain only English texts and focus on puns.
Approach: They collected a dataset of jokes and funny dialogues in Russian and complemented them carefully with unfunny texts with similar lexical properties.
Outcome: The proposed method is based on the universal language model finetuning and has an F1 score of 0.91 on a test set.
HAHA 2019 Dataset: A Corpus for Humor Analysis in Spanish (2020.lrec-1)

Copied to clipboard

Challenge: 30,000 Spanish tweets were crowd-annotated with humor value and funniness score . the corpus contains approximately 38.6% of humorous tweets with an average score of 2.04 in a scale from 1 to 5 for the humorous tweet.
Approach: They develop a corpus of 30,000 Spanish tweets crowd-annotated with humor value and funniness score.
Outcome: The results obtained from the 30,000 tweets in the Spanish language are encouraging.
Building a Sentiment Corpus of Tweets in Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: Sentiment analysis is a popular area of Natural Language Processing due to its subjective and semantic characteristics.
Approach: They propose to annotate Brazilian Portuguese sentences manually using a sentiment corpus . they run experiments on polarity classification using six machine learning classifiers .
Outcome: The proposed method is based on a Brazilian Portuguese sentiment corpus and achieved 80.38% on F-Measure and 64.87% when including the neutral class.
AIA-BDE: A Corpus of FAQs in Portuguese and their Variations (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of 380 domain-oriented FAQs in Portuguese is presented . paraphrases or entailed questions are created manually, by humans, or automatically, with Google Translate.
Approach: They present a corpus of 380 domain-oriented FAQs in Portuguese and their variations, i.e., paraphrases or entailed questions, created manually, by humans, or automatically, with Google Translate.
Outcome: The proposed system outperforms other systems in the domain of question-answering . it performs well when matching variations with their original questions .
A Survey of Pun Generation: Datasets, Evaluations and Methodologies (2025.findings-emnlp)

Copied to clipboard

Challenge: Pun generation aims to modify linguistic elements in text to produce humour or evoke double meanings.
Approach: They propose to review pun generation datasets and methods across different stages . pun generation aims to produce humour or evoke double meanings .
Outcome: This paper summarises both automated and human evaluation metrics used to assess the quality of pun generation.
Development and Validation of a Corpus for Machine Humor Comprehension (2020.lrec-1)

Copied to clipboard

Challenge: a Chinese humor corpus was labeled with five levels of funniness, eight skill sets of humor, and six dimensions of intent by only one annotator.
Approach: They develop a Chinese humor corpus with 3,365 jokes labeled with five levels of funniness, eight skill sets of humor, and six dimensions of intent by only one annotator.
Outcome: The proposed corpus contains 3,365 jokes from over 40 sources.
PPORTAL_ner: An Annotated Corpus of Portuguese Literary Entities (2024.lrec-main)

Copied to clipboard

Challenge: Annotated corpus of 25 literary texts provides a rich set of annotations for Named Entity Recognition models.
Approach: They propose an annotation dataset that simplifies the development of Named Entity Recognition models for Portuguese literary texts.
Outcome: The proposed dataset simplifies the development of Named Entity Recognition models for Portuguese literary works.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations