Obtaining Reliable Human Ratings of Valence, Arousal, and Dominance for 20,000 English Words (P18-1)
Copied to clipboard
| Challenge: | Words play a central role in language and thought. |
| Approach: | They propose a Lexicon with ratings of valence, arousal, and dominance for 20,000 words . they use Best–Worst Scaling to obtain fine-grained scores . |
| Outcome: | The proposed Lexicon has human ratings of valence, arousal, and dominance for 20,000 words . the ratings are more reliable than those in existing lexicons, the authors show . |
Similar Papers
Word Affect Intensities (L18-1)
Copied to clipboard
| Challenge: | Existing lexicons of affect only show coarse associations, but are not accurate as human-created ones. |
| Approach: | They propose to use a manually created affect intensity lexicon with real-valued intensity scores for anger, fear, joy, and sadness. |
| Outcome: | The lexicon has real-valued scores for anger, fear, joy, and sadness . anger, fears, and sad words have very similar VAD scores . |
ValNorm Quantifies Semantics to Reveal Consistent Valence Biases Across Languages and Over Centuries (2021.emnlp-main)
Copied to clipboard
| Challenge: | Word embeddings learn implicit biases from word co-occurrence statistics . valNorm is a new intrinsic evaluation task and method to quantify affect in word embedded word sets . |
| Approach: | They propose a method to quantify valence dimension of affect in human-rated word sets . they apply ValNorm to embeddings from seven languages and 200 years of text . |
| Outcome: | The proposed method achieves a high accuracy in quantifying the valence of non-discriminatory, non-social group word sets. |
Guilt by Association: Emotion Intensities in Lexical Representations (2021.emnlp-main)
Copied to clipboard
| Challenge: | linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks . |
| Approach: | They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models. |
| Outcome: | The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data. |
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)
Copied to clipboard
| Challenge: | a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes . |
| Approach: | They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones . |
| Outcome: | The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions. |
Words of Warmth: Trust and Sociability Norms for over 26k English Words (2025.acl-long)
Copied to clipboard
| Challenge: | Social psychologists have shown that Warmth (W) and Competence (C) are the primary dimensions along which we assess other people and groups. |
| Approach: | They propose a repository of word–warmth and word–trust associations for over 26k English words. |
| Outcome: | The proposed lexicon enables bias and stereotype research through case studies on target entities. |
Scalar Adjective Identification and Multilingual Ranking (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing studies on scalar adjective ranking have focused on English due to the availability of datasets for evaluation. |
| Approach: | They propose a binary classification task to examine the models’ ability to distinguish scalar from relational adjectives in English. |
| Outcome: | The proposed task compares the models' ability to distinguish scalar from relational adjectives in English using monolingual and multilingual models. |
Ruddit: Norms of Offensiveness for English Reddit Comments (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods to detect offensive language have been limited by categorical labels . however, there are several challenges in the detection of such content . |
| Approach: | They analyze Reddit comments with fine-grained, real-valued offensiveness scores . they evaluate the ability of widely-used neural models to predict offensiveness . |
| Outcome: | The proposed method produces highly reliable offensiveness scores and can predict scores on reddit comments. |
A Human Evaluation of AMR-to-English Generation Systems (2020.coling-main)
Copied to clipboard
| Challenge: | a recent human evaluation of AMR generation systems is compared to automated metrics. |
| Approach: | They propose a human evaluation which collects fluency and adequacy scores and categorization of error types for AMR generation systems. |
| Outcome: | The results show that human evaluations are more nuanced than automated metrics. |
Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 3 (Industry Papers) (N18-3)
Copied to clipboard
| Challenge: | NAACL 2018 Industry Track aims to provide a forum for researchers, engineers and application developers to share their experience in real-world language problems. |
| Approach: | NAACL 2018 Industry Track is the inaugural conference in the *ACL family of conferences . organizers wanted to provide a forum for researchers, engineers and application developers to share their experience . six of the papers were desk rejects due to non-conformance with submission requirements . |
| Outcome: | the inaugural industry track at NAACL 2018 received 91 submissions, exceeding expectations . the track will focus on problems that manifest themselves more readily in industry . |
“Fifty Shades of Bias”: Normative Ratings of Gender Bias in GPT Generated English Text (2023.emnlp-main)
Copied to clipboard
| Challenge: | Prior work treats gender bias as a binary classification task, but a comparative annotation framework can be used to assess the impact of biases. |
| Approach: | They propose to generate a dataset with normative ratings of gender bias in English text with a comparative annotation framework. |
| Outcome: | The first dataset of GPT-generated English text with normative ratings of gender bias is analyzed using Best–Worst Scaling . |