User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)

Copied to clipboard

Challenge: Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias.
Approach: They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC.
Outcome: The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community.

Similar Papers

The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)

Copied to clipboard

Challenge: Social media data is often aggregated without regard to users in the Twitter populations of each community.
Approach: They propose to use Twitter language to build community-level models using Twitter language aggregated by users.
Outcome: The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets.
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans .
Approach: They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs.
Outcome: The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs.
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models.
Approach: They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups.
Outcome: The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech.
Incorporating Textual Information on User Behavior for Personality Prediction (P19-2)

Copied to clipboard

Challenge: Recent studies have shown that textual information of user posts and user behaviors are useful for predicting the personality of social media users.
Approach: They propose to use textual information of user behaviors to predict personality of Twitter users by taking user behaviors into account.
Outcome: The proposed models can predict personality of users who do not post frequently, while taking user behaviors into account.
Towards Open-Domain Twitter User Profile Inference (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to user profile inference focus on limited attributes and can reveal users' private information.
Approach: They propose a prompt-based generation method which can infer values implicitly mentioned in Twitter user profiles.
Outcome: The proposed method can infer more comprehensive user profiles than baseline extraction-based methods, but limitations remain to be applied for real-world use.
Hierarchical Modeling for User Personality Prediction: The Role of Message-Level Attention (2020.acl-main)

Copied to clipboard

Challenge: Language processing is increasingly finding use as a supplement for questionnaires to assess psychological attributes of consenting individuals, but most approaches neglect to consider whether all documents of an individual are equally informative.
Approach: They propose a model that uses message-level attention to learn the relative weight of users’ social media posts for assessing their five factor personality traits.
Outcome: The proposed model outperforms models with word-level attention and yields state-of-the-art accuracies for all five personality traits.
Multi-Modal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision–Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Recent advances in self-supervised training have led to a new class of pretrained vision–language models.
Approach: They propose a visual and textual bias benchmark to assess bias in self-supervised multimodal models using 3,800 images and phrases from 14 population subgroups.
Outcome: The proposed model shows that it favors certain groups while maintaining the accuracy of the model.
Nationality Bias in Text Generation (2023.eacl-main)

Copied to clipboard

Challenge: Existing studies have shown that nationality biases in language models can be a factor in improving the performance of social NLP models.
Approach: They propose to use a text generation model, GPT-2, to analyze how the number of internet users and the country’s economic status affects the sentiment of stories.
Outcome: The proposed model accentuates biases about country-based demonyms and reduces them with the use of adversarial triggering.
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes.
Approach: They assemble and publish a multilingual Twitter corpus for the task of hate speech detection using inferred author demographic factors.
Outcome: The results show that the classifiers learn human biases and can be discriminatory towards certain demographic groups.
The Face of Persuasion: Analyzing Bias and Generating Culture-Aware Ads (2025.findings-emnlp)

Copied to clipboard

Challenge: Text-to-image models are appealing for customizing visual ads and targeting specific populations.
Approach: We examine the disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed.
Outcome: The proposed technique is based on a demographic bias analysis of ads for different topics and a disparate level of persuasiveness of ads that are identical except for gender/race of the people portrayed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations