Papers by Daniel Preoţiuc-Pietro

8 papers
The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)

Copied to clipboard

Challenge: Social media data is often aggregated without regard to users in the Twitter populations of each community.
Approach: They propose to use Twitter language to build community-level models using Twitter language aggregated by users.
Outcome: The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets.
Why Swear? Analyzing and Inferring the Intentions of Vulgar Expressions (D18-1)

Copied to clipboard

Challenge: Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication.
Approach: They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories.
Outcome: The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes.
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)

Copied to clipboard

Challenge: Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias.
Approach: They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC.
Outcome: The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community.
Multi-task Pairwise Neural Ranking for Hashtag Segmentation (P19-1)

Copied to clipboard

Challenge: Hashtags are used to add metadata to textual utterances, but their semantic content is difficult to infer as they often contain multiple tokens joined together.
Approach: They propose to use a dataset of 12,594 hashtags to infer hashtag semantics . they propose to frame the problem as a pairwise ranking problem between candidate segmentations .
Outcome: The proposed methods show 24.6% error reduction in hashtag segmentation accuracy compared to the current state-of-the-art method.
Categorizing and Inferring the Relationship between the Text and Image of Twitter Posts (P19-1)

Copied to clipboard

Challenge: Social media posts often contain images to provide content, provide context, or express feelings.
Approach: They build and release a dataset of image tweets annotated with four different classes which express whether the text or the image provides additional information to the other modality.
Outcome: The proposed method can be used in several downstream applications including pre-training image tagging models and collecting distantly supervised data for image captioning.
Automatically Identifying Complaints in Social Media (P19-1)

Copied to clipboard

Challenge: Complaining is a basic speech act used to express a negative mismatch between reality and expectations in a particular situation.
Approach: They present a systematic analysis of complaints in computational linguistics . they collect annotated data set of written complaints expressed on Twitter .
Outcome: The proposed model achieves predictive performance of up to 79 F1 using distant supervision.
Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)

Copied to clipboard

Challenge: Vulgarity is a common linguistic expression and is used to perform several linguistic functions.
Approach: They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance.
Outcome: The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings.
Analyzing Linguistic Differences between Owner and Staff Attributed Tweets (P19-1)

Copied to clipboard

Challenge: Existing studies on social media assume that all tweets are authored by the same person . this paper examines the linguistic differences between posts signed by the account owner and staff attributed to the account's owner .
Approach: They analyze linguistic differences between tweets signed by the account owner and staff attributed to their staff.
Outcome: The proposed model predicts owner and staff attributed tweets with good accuracy even without training data from the account.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations