| Challenge: | Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias. |
| Approach: | They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC. |
| Outcome: | The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community. |
Similar Papers
The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)
Copied to clipboard
Salvatore Giorgi, Daniel Preoţiuc-Pietro, Anneke Buffone, Daniel Rieman, Lyle Ungar, H. Andrew Schwartz
| Challenge: | Social media data is often aggregated without regard to users in the Twitter populations of each community. |
| Approach: | They propose to use Twitter language to build community-level models using Twitter language aggregated by users. |
| Outcome: | The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets. |
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans . |
| Approach: | They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs. |
| Outcome: | The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs. |
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models. |
| Approach: | They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups. |
| Outcome: | The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech. |
Incorporating Textual Information on User Behavior for Personality Prediction (P19-2)
Copied to clipboard
| Challenge: | Recent studies have shown that textual information of user posts and user behaviors are useful for predicting the personality of social media users. |
| Approach: | They propose to use textual information of user behaviors to predict personality of Twitter users by taking user behaviors into account. |
| Outcome: | The proposed models can predict personality of users who do not post frequently, while taking user behaviors into account. |
Towards Open-Domain Twitter User Profile Inference (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to user profile inference focus on limited attributes and can reveal users' private information. |
| Approach: | They propose a prompt-based generation method which can infer values implicitly mentioned in Twitter user profiles. |
| Outcome: | The proposed method can infer more comprehensive user profiles than baseline extraction-based methods, but limitations remain to be applied for real-world use. |
Hierarchical Modeling for User Personality Prediction: The Role of Message-Level Attention (2020.acl-main)
Copied to clipboard
| Challenge: | Language processing is increasingly finding use as a supplement for questionnaires to assess psychological attributes of consenting individuals, but most approaches neglect to consider whether all documents of an individual are equally informative. |
| Approach: | They propose a model that uses message-level attention to learn the relative weight of users’ social media posts for assessing their five factor personality traits. |
| Outcome: | The proposed model outperforms models with word-level attention and yields state-of-the-art accuracies for all five personality traits. |
Multi-Modal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision–Language Models (2023.eacl-main)
Copied to clipboard
| Challenge: | Recent advances in self-supervised training have led to a new class of pretrained vision–language models. |
| Approach: | They propose a visual and textual bias benchmark to assess bias in self-supervised multimodal models using 3,800 images and phrases from 14 population subgroups. |
| Outcome: | The proposed model shows that it favors certain groups while maintaining the accuracy of the model. |
Nationality Bias in Text Generation (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing studies have shown that nationality biases in language models can be a factor in improving the performance of social NLP models. |
| Approach: | They propose to use a text generation model, GPT-2, to analyze how the number of internet users and the country’s economic status affects the sentiment of stories. |
| Outcome: | The proposed model accentuates biases about country-based demonyms and reduces them with the use of adversarial triggering. |
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. |
| Approach: | They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors. |
| Outcome: | The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups. |
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Text-to-image models are appealing for customizing visual ads and targeting specific populations. |
| Approach: | We examine the disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed. |
| Outcome: | The proposed technique is based on a demographic bias analysis of ads for different topics and a disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed. |