Papers by Daniel Preoţiuc-Pietro
The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)
Copied to clipboard
Salvatore Giorgi, Daniel Preoţiuc-Pietro, Anneke Buffone, Daniel Rieman, Lyle Ungar, H. Andrew Schwartz
| Challenge: | Social media data is often aggregated without regard to users in the Twitter populations of each community. |
| Approach: | They propose to use Twitter language to build community-level models using Twitter language aggregated by users. |
| Outcome: | The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets. |
Why Swear? Analyzing and Inferring the Intentions of Vulgar Expressions (D18-1)
Copied to clipboard
| Challenge: | Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication. |
| Approach: | They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories. |
| Outcome: | The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes. |
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)
Copied to clipboard
| Challenge: | Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias. |
| Approach: | They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC. |
| Outcome: | The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community. |
Multi-task Pairwise Neural Ranking for Hashtag Segmentation (P19-1)
Copied to clipboard
| Challenge: | Hashtags are used to add metadata to textual utterances, but their semantic content is difficult to infer as they often contain multiple tokens joined together. |
| Approach: | They propose to use a dataset of 12,594 hashtags to infer hashtag semantics . they propose to frame the problem as a pairwise ranking problem between candidate segmentations . |
| Outcome: | The proposed methods show 24.6% error reduction in hashtag segmentation accuracy compared to the current state-of-the-art method. |
Categorizing and Inferring the Relationship between the Text and Image of Twitter Posts (P19-1)
Copied to clipboard
| Challenge: | Social media posts often contain images to provide content, provide context, or express feelings. |
| Approach: | They build and release a dataset of image tweets annotated with four different classes which express whether the text or the image provides additional information to the other modality. |
| Outcome: | The proposed method can be used in several downstream applications including pre-training image tagging models and collecting distantly supervised data for image captioning. |
Automatically Identifying Complaints in Social Media (P19-1)
Copied to clipboard
| Challenge: | Complaining is a basic speech act used to express a negative mismatch between reality and expectations in a particular situation. |
| Approach: | They present a systematic analysis of complaints in computational linguistics . they collect annotated data set of written complaints expressed on Twitter . |
| Outcome: | The proposed model achieves predictive performance of up to 79 F1 using distant supervision. |
Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)
Copied to clipboard
| Challenge: | Vulgarity is a common linguistic expression and is used to perform several linguistic functions. |
| Approach: | They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance. |
| Outcome: | The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings. |
Analyzing Linguistic Differences between Owner and Staff Attributed Tweets (P19-1)
Copied to clipboard
| Challenge: | Existing studies on social media assume that all tweets are authored by the same person . this paper examines the linguistic differences between posts signed by the account owner and staff attributed to the account's owner . |
| Approach: | They analyze linguistic differences between tweets signed by the account owner and staff attributed to their staff. |
| Outcome: | The proposed model predicts owner and staff attributed tweets with good accuracy even without training data from the account. |