HappyDB: A Corpus of 100,000 Crowdsourced Happy Moments (L18-1)

Copied to clipboard

Challenge: Recent research has focused on developing technologies that help users incorporate the findings of the science of happiness into their daily lives.
Approach: They crowd-sourced HappyDB, a corpus of 100,000 happy moments, and applied several state-of-the-art analysis techniques to analyze HappyDB.
Outcome: The proposed technology can understand how people express their happy moments in text and analyze them using state-of-the-art techniques.

Similar Papers

GoodNewsEveryone: A Corpus of News Headlines Annotated with Emotions, Semantic Roles, and Reader Perception (2020.lrec-1)

Copied to clipboard

Challenge: Fewer studies address emotions as a phenomenon to be tackled with structured learning, which can be explained by the lack of relevant datasets.
Approach: They propose to annotate 5000 English news headlines with their associated emotions, the corresponding emotion experiencers and textual cues, related emotion causes and targets, and the reader’s perception of the emotion of the headline.
Outcome: The proposed method enables further research on emotion classification, emotion intensity prediction, emotion cause detection and supports qualitative studies.
An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing frameworks for emotion recognition are limited and do not allow for categorical versus dimensional oppositions.
Approach: They propose to use the emotions joy, love, anger, sadness and fear as well as dimensional models to annotate texts from different domains and topics.
Outcome: The proposed frameworks are well-suited to annotate texts from different domains and topics, but the connotation of the labels strongly depends on the origin of the texts.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Proceedings of the First Workshop on Aggregating and Analysing Crowdsourced Annotations for NLP (D19-59)

Copied to clipboard

Challenge: The first workshop on crowdsourcing for NLP is open to all .
Approach: The first workshop on crowdsourcing annotations for NLP is held at the acl.com . the workshop will focus on methods for aggregating and analysing crowdsourced data for Nl-specific tasks.
Outcome: The first workshop on crowdsourcing for NLP received 16 submissions and accepted 7 . the workshop will focus on ambiguous, subjective or ambiguity analysis of crowdsourced data .
Construction of the Corpus of Everyday Japanese Conversation: An Interim Report (L18-1)

Copied to clipboard

Challenge: a new corpus of everyday conversations is being developed in the field of everyday conversation . the corpus is based on 94 hours of recordings of everyday Japanese conversations .
Approach: They propose to build a large-scale corpus of everyday Japanese conversation in a balanced manner.
Outcome: The proposed corpus will be published in 2022 and consist of more than 200 hours of recordings.
A Dataset of Crowdsourced Word Sequences: Collections and Answer Aggregation for Ground Truth Creation (D19-59)

Copied to clipboard

Challenge: Existing work on answer aggregation for labels is limited . existing work on label aggregations is limited to label .
Approach: They propose three approaches to extractive word sequence aggregation from translated sentences generated by multiple workers.
Outcome: The proposed dataset contains translated sentences generated from multiple workers.
Event2Mind: Commonsense Inference on Events, Intents, and Reactions (P18-1)

Copied to clipboard

Challenge: Using a crowdsourced corpus of 25,000 event phrases, we construct a new task that uses commonsense reasoning to reason about the likely intents and reactions of the event participants.
Approach: They construct a crowdsourced corpus of 25,000 event phrases and use them to construct 'commonsense inference' they demonstrate that neural encoder-decoder models can compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants.
Outcome: The proposed task can be used to uncover implicit gender inequality in movie scripts.
Love Me, Love Me, Say (and Write!) that You Love Me: Enriching the WASABI Song Corpus with Lyrics Annotations (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of songs enriched with metadata extracted from music databases on the Web contains 1.73M songs with lyrics (1.41M unique lyrics) a researcher proposes methods to extract relevant information from lyrics, including their structure segmentation, topic, explicitness of lyrics content, salient passages of a song and emotions conveyed.
Approach: They propose to extract relevant information from lyrics by using music databases . they propose to use metadata extracted from music databases to analyze lyrics .
Outcome: The proposed methods can be exploited by music search engines and music professionals to better handle large collections of lyrics.
EMTC: Multilabel Corpus in Movie Domain for Emotion Analysis in Conversational Text (L18-1)

Copied to clipboard

Challenge: Existing emotion corpora collected from twitters and use hashtags are limited in the number of characters.
Approach: They propose to build an emotion corpus based on conversational text data that includes 2.1 million utterances and is partly annotated by ourselves and independent annotators.
Outcome: The proposed corpus includes conversations from movies with more than 2.1 million utterances which are partly annotated by ourselves and independent annotators.
EmotionLines: An Emotion Corpus of Multi-Party Conversations (L18-1)

Copied to clipboard

Challenge: Emotion is a critical characteristic to distinguish people from machines.
Approach: They propose a dataset with emotions labeling on all utterances in each dialogue . they use Friends TV scripts and Facebook messenger dialogues to collect the data .
Outcome: The proposed dataset is the first with emotions labeling on all utterances in each dialogue based on their textual content.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations