| Challenge: | Existing methods for analyzing people portrayals take an unsupervised approach, or rely on domain-specific knowledge. |
| Approach: | They show how contextualized word embeddings can be used to capture affect dimensions in portrayals of people. |
| Outcome: | The proposed method can capture affect dimensions in portrayals of men and women . it is biased towards training data, which limits its usefulness to in-domain analyses . |
Similar Papers
Interpreting Pretrained Contextualized Representations via Reductions to Static Embeddings (2020.acl-main)
Copied to clipboard
| Challenge: | Contextualized representations have become the default for downstream NLP applications. |
| Approach: | They propose a method for converting from contextualized representations to static lookup-table embeddings and apply it to 5 popular pretrained models and 9 sets of pretrained weights. |
| Outcome: | The proposed methods show that pooling over many contexts significantly improves representational quality under intrinsic evaluation. |
Gender Bias in Contextualized Word Embeddings (N19-1)
Copied to clipboard
| Challenge: | Existing studies show that training word embeddings in large corpora could lead to encoding societal biases present in these human-produced data. |
| Approach: | They conduct several intrinsic analyses to quantify, analyze and mitigate gender bias exhibited in ELMo’s contextualized word vectors. |
| Outcome: | The proposed method mitigates gender bias on WinoBias probing corpus and demonstrates that it can be implemented in other systems. |
How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings (D19-1)
Copied to clipboard
| Challenge: | Existing word embeddings were static, requiring all senses of a polysemous word to share the same representation. |
| Approach: | They found that the contextualized representations of all words are not isotropic in any layer of the contextualizing model. |
| Outcome: | The results show that the representations of all words are not isotropic in any layer of the contextualizing model. |
Unequal Representations: Analyzing Intersectional Biases in Word Embeddings Using Representational Similarity Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Specifically, we probe contextualized and non-contextualized word embeddings for evidence of intersectional biases against Black women. |
| Approach: | They propose a representational similarity analysis approach to detect human-like biases in word embeddings using representational similarities analysis. |
| Outcome: | The proposed approach aligns with intersectionality theory, which states that multiple identity categories layer on top of each other to create unique modes of discrimination that are not shared by any individual category. |
Building Static Embeddings from Contextual Ones: Is It Useful for Building Distributional Thesauri? (2022.lrec-1)
Copied to clipboard
| Challenge: | contextual language models are dominant in the field of Natural Language Processing, but they are not suitable for all uses. |
| Approach: | They propose a method for building word or type-level embeddings from contextual models . they evaluate a large set of English nouns from the perspective of extracting semantic similarity relations . |
| Outcome: | The proposed method can be used to build word or type embeddings from contextual models . it can be exploited for a wide set of English nouns, showing it can improve distributional thesauri . |
Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding Words (2020.acl-main)
Copied to clipboard
| Challenge: | a suite of probing tasks test contextual embeddings for encoding of information about surrounding words . authors: little is known about what information embeddables encode about the context words encode . a recent study shows that contextual embeds can be powerful for many tasks . |
| Approach: | They propose probing tasks that enable fine-grained testing of contextual embeddings . they examine popular contextual encoders and find that each encodes contextual information across tokens a little different . |
| Outcome: | The proposed probing tasks show that word embeddings encode information about words . the tests show that the encoded information is encoded across tokens with near-perfect recoverability . |
Searching for the X-Factor: Exploring Corpus Subjectivity for Word Embeddings (P18-1)
Copied to clipboard
| Challenge: | Existing word embedding methods for natural language processing are limited in their ability to produce dense word embeds. |
| Approach: | They propose a word embedding SentiVec which is infused with sentiment information from a lexical resource and outperforms baselines on subjectivity-sensitive tasks. |
| Outcome: | The proposed word embedding SentiVec outperforms baselines on subjectivity-sensitive tasks. |
When do Word Embeddings Accurately Reflect Surveys on our Beliefs About People? (2020.acl-main)
Copied to clipboard
| Challenge: | a study of word embeddings shows that social biases are more accurate than survey data for some dimensions of meaning. |
| Approach: | a new study investigates the extent to which word embeddings accurately reflect biases . they find that biased word embeds mirror survey data across 17 dimensions of social meaning . |
| Outcome: | a new study shows that word embeddings accurately reflect biases on average across dimensions of social meaning . biased embedders are more reflective of survey data for some dimensions of meaning than others, the study finds . |
Measuring Context-Word Biases in Lexical Semantic Datasets (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing pretrained contextualized models have been used to evaluate word-in-context representations in many lexical semantic tasks. |
| Approach: | They propose to quantify the degree of context or word biases in existing datasets by probing masked input. |
| Outcome: | The proposed model performs better when both word and context are available than with masked input. |
Quantifying the Contextualization of Word Representations with Semantic Class Probing (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Pretrained language models are effective in solving NLP tasks, but there are still questions about how and why they work so well. |
| Approach: | They use BERT to quantify contextualization by studying the extent of inference . they show that top layer representations support highly accurate inference of semantic classes . |
| Outcome: | The proposed model is highly accurate, but weak in the lower layers . it is more task-specific after finetuning while lower layers are more transferable . |