Challenge: Existing approaches to predicting Twitter users' demographic attributes exploit, select, and combine various features generated from text and network to achieve the best performance.
Approach: They extend existing Twitter occupational class prediction data set and exploit social network homophily to achieve competitive performance.
Outcome: The proposed method achieves better performance on a dataset with a small fraction of the training data.

Similar Papers

Cross-media User Profiling with Joint Textual and Social User Embedding (C18-1)

Copied to clipboard

Challenge: Empirical studies demonstrate the effectiveness of the proposed approach to cross-media user profiling tasks.
Approach: They propose a uniform user embedding learning approach to address cross-media user profiling by bridging the knowledge between the source and target media.
Outcome: Empirical results show that the proposed approach performs well on two cross-media user profiling tasks.
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)

Copied to clipboard

Challenge: Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias.
Approach: They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC.
Outcome: The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community.
You Shall Know a User by the Company It Keeps: Dynamic Representations for Social Media Users in NLP (D19-1)

Copied to clipboard

Challenge: Current approaches to social media modelling ignore the fact that an individual may be part of several communities which are not equally relevant in all communicative situations.
Approach: They propose a model that captures the sociological phenomenon of homophily and combines it with linguistic information to make a prediction.
Outcome: The proposed model significantly outperforms existing models on three different tasks and is compared with other models.
Automatic Classification of Students on Twitter Using Simple Profile Information (2020.aacl-srw)

Copied to clipboard

Challenge: Existing models for age classification of students and non-students are restrictive and require access to many tweets.
Approach: They propose a model which uses 3 tweet-content features to classify users as students or non-students.
Outcome: The proposed model achieves an accuracy of 88.1% and an F1 score of .704 compared to previous models, which require access to many user tweets.
Label Embedding using Hierarchical Structure of Labels for Twitter Classification (D19-1)

Copied to clipboard

Challenge: Twitter is used for disaster monitoring and news material gathering . we propose a method that can consider the hierarchical structure of labels and labels themselves .
Approach: They propose a method that can consider the hierarchical structure of labels and label texts themselves.
Outcome: The proposed method outperforms the methods of the conference participants over the text REtrieval Conference (TREC) 2018 Incident Streams (IS) dataset.
A Hierarchical Location Prediction Neural Network for Twitter User Geolocation (D19-1)

Copied to clipboard

Challenge: Existing methods to estimate user location ignore hierarchical structure among locations.
Approach: They propose a hierarchical location prediction neural network for Twitter user geolocation that first predicts the home country for a user, then uses the country result to guide the city-level prediction.
Outcome: The proposed model can achieve state-of-the-art results over three common benchmarks under different feature settings and greatly reduces the mean error distance.
Predicting Human Activities from User-Generated Content (P19-1)

Copied to clipboard

Challenge: Several studies have applied computational approaches to the understanding and modeling of human behavior at scale and in real time.
Approach: They propose a sentence embedding framework tailored to recognize the semantics of human activities and perform automatic clustering of these activities.
Outcome: The proposed framework can make predictions based on the text of user-generated content and self-description.
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models.
Approach: They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups.
Outcome: The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech.
Incorporating Textual Information on User Behavior for Personality Prediction (P19-2)

Copied to clipboard

Challenge: Recent studies have shown that textual information of user posts and user behaviors are useful for predicting the personality of social media users.
Approach: They propose to use textual information of user behaviors to predict personality of Twitter users by taking user behaviors into account.
Outcome: The proposed models can predict personality of users who do not post frequently, while taking user behaviors into account.
The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)

Copied to clipboard

Challenge: Social media data is often aggregated without regard to users in the Twitter populations of each community.
Approach: They propose to use Twitter language to build community-level models using Twitter language aggregated by users.
Outcome: The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations