| Challenge: | Existing approaches to predicting Twitter users' demographic attributes exploit, select, and combine various features generated from text and network to achieve the best performance. |
| Approach: | They extend existing Twitter occupational class prediction data set and exploit social network homophily to achieve competitive performance. |
| Outcome: | The proposed method achieves better performance on a dataset with a small fraction of the training data. |
Similar Papers
Cross-media User Profiling with Joint Textual and Social User Embedding (C18-1)
Copied to clipboard
| Challenge: | Empirical studies demonstrate the effectiveness of the proposed approach to cross-media user profiling tasks. |
| Approach: | They propose a uniform user embedding learning approach to address cross-media user profiling by bridging the knowledge between the source and target media. |
| Outcome: | Empirical results show that the proposed approach performs well on two cross-media user profiling tasks. |
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)
Copied to clipboard
| Challenge: | Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias. |
| Approach: | They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC. |
| Outcome: | The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community. |
You Shall Know a User by the Company It Keeps: Dynamic Representations for Social Media Users in NLP (D19-1)
Copied to clipboard
| Challenge: | Current approaches to social media modelling ignore the fact that an individual may be part of several communities which are not equally relevant in all communicative situations. |
| Approach: | They propose a model that captures the sociological phenomenon of homophily and combines it with linguistic information to make a prediction. |
| Outcome: | The proposed model significantly outperforms existing models on three different tasks and is compared with other models. |
Automatic Classification of Students on Twitter Using Simple Profile Information (2020.aacl-srw)
Copied to clipboard
| Challenge: | Existing models for age classification of students and non-students are restrictive and require access to many tweets. |
| Approach: | They propose a model which uses 3 tweet-content features to classify users as students or non-students. |
| Outcome: | The proposed model achieves an accuracy of 88.1% and an F1 score of .704 compared to previous models, which require access to many user tweets. |
Label Embedding using Hierarchical Structure of Labels for Twitter Classification (D19-1)
Copied to clipboard
| Challenge: | Twitter is used for disaster monitoring and news material gathering . we propose a method that can consider the hierarchical structure of labels and labels themselves . |
| Approach: | They propose a method that can consider the hierarchical structure of labels and label texts themselves. |
| Outcome: | The proposed method outperforms the methods of the conference participants over the text REtrieval Conference (TREC) 2018 Incident Streams (IS) dataset. |
A Hierarchical Location Prediction Neural Network for Twitter User Geolocation (D19-1)
Copied to clipboard
| Challenge: | Existing methods to estimate user location ignore hierarchical structure among locations. |
| Approach: | They propose a hierarchical location prediction neural network for Twitter user geolocation that first predicts the home country for a user, then uses the country result to guide the city-level prediction. |
| Outcome: | The proposed model can achieve state-of-the-art results over three common benchmarks under different feature settings and greatly reduces the mean error distance. |
Predicting Human Activities from User-Generated Content (P19-1)
Copied to clipboard
| Challenge: | Several studies have applied computational approaches to the understanding and modeling of human behavior at scale and in real time. |
| Approach: | They propose a sentence embedding framework tailored to recognize the semantics of human activities and perform automatic clustering of these activities. |
| Outcome: | The proposed framework can make predictions based on the text of user-generated content and self-description. |
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models. |
| Approach: | They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups. |
| Outcome: | The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech. |
Incorporating Textual Information on User Behavior for Personality Prediction (P19-2)
Copied to clipboard
| Challenge: | Recent studies have shown that textual information of user posts and user behaviors are useful for predicting the personality of social media users. |
| Approach: | They propose to use textual information of user behaviors to predict personality of Twitter users by taking user behaviors into account. |
| Outcome: | The proposed models can predict personality of users who do not post frequently, while taking user behaviors into account. |
The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)
Copied to clipboard
Salvatore Giorgi, Daniel Preoţiuc-Pietro, Anneke Buffone, Daniel Rieman, Lyle Ungar, H. Andrew Schwartz
| Challenge: | Social media data is often aggregated without regard to users in the Twitter populations of each community. |
| Approach: | They propose to use Twitter language to build community-level models using Twitter language aggregated by users. |
| Outcome: | The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets. |