| Challenge: | Existing methods for learning to compare social media users fail to generalize to new users or even to previously known users. |
| Approach: | They propose a procedure to learn a mapping from short episodes of user activity to a vector space in which the distance between points captures the similarity of the corresponding users’ invariant features. |
| Outcome: | The proposed procedure can be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space. |
Similar Papers
A Deep Metric Learning Approach to Account Linking (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to identify abusive content may fail to adapt to new trends, and individual posts may fail . |
| Approach: | They propose a method that embeds variable-sized samples of user activity into a vector space, where samples by the same author map to nearby points. |
| Outcome: | The proposed model outperforms several competitive baselines under a new evaluation framework modeled after established benchmarks in other domains. |
You Shall Know a User by the Company It Keeps: Dynamic Representations for Social Media Users in NLP (D19-1)
Copied to clipboard
| Challenge: | Current approaches to social media modelling ignore the fact that an individual may be part of several communities which are not equally relevant in all communicative situations. |
| Approach: | They propose a model that captures the sociological phenomenon of homophily and combines it with linguistic information to make a prediction. |
| Outcome: | The proposed model significantly outperforms existing models on three different tasks and is compared with other models. |
Simple Attention-Based Representation Learning for Ranking Short Social Media Posts (N19-1)
Copied to clipboard
| Challenge: | Existing approaches to ranking short social media posts are complex and require different components to capture a multitude of relevance signals. |
| Approach: | They propose a word-level Siamese architecture with attention-based mechanisms for capturing semantic "soft" matches between query and post tokens. |
| Outcome: | The proposed model is faster and simpler than existing models and more efficient than existing approaches. |
PASUM: A Pre-training Architecture for Social Media User Modeling Based on Text Graph (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have incorporated different digital traces to better learn the representations of social media users, limited by overloaded text information and hard-to-collect social network information. |
| Approach: | They propose a Pre-training Architecture for Social Media User Modeling based on Text Graph and combine microblogs to represent social media users based upon the text graph model. |
| Outcome: | The proposed framework can represent users based on text even without social network information on microblogs. |
Creation and evaluation of timelines for longitudinal user posts (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for segmenting user posts into timelines improve quality and cost of manual annotation. |
| Approach: | They propose a set of methods for segmenting longitudinal user posts into timelines likely to contain interesting moments of change in a user’s behaviour based on their online posting activity. |
| Outcome: | The proposed framework is able to evaluate two different social media datasets and compares with existing models. |
Cross-media User Profiling with Joint Textual and Social User Embedding (C18-1)
Copied to clipboard
| Challenge: | Empirical studies demonstrate the effectiveness of the proposed approach to cross-media user profiling tasks. |
| Approach: | They propose a uniform user embedding learning approach to address cross-media user profiling by bridging the knowledge between the source and target media. |
| Outcome: | Empirical results show that the proposed approach performs well on two cross-media user profiling tasks. |
Extracting Age-Related Stereotypes from Social Media Texts (2022.lrec-1)
Copied to clipboard
| Challenge: | a method for extracting age-related stereotypes from Twitter data is under-studied in NLP . stereotyping on the basis of protected characteristics has been understudied . |
| Approach: | They propose a method for extracting age-related stereotypes from Twitter data . they generate a corpus of 300,000 over-generalizations about four contemporary generations . |
| Outcome: | The method uncovers common stereotypes as reported in media and psychological literature . it also finds that stereotypes for different generations vary across topics . |
Adapting Deep Learning Methods for Mental Health Prediction on Social Media (D19-55)
Copied to clipboard
| Challenge: | a quarter of the population in Europe suffers from an episode of a mental disorder in their life, according to the World Health Organization . text analysis of rich resources like social media can contribute to deeper understanding of mental health and provide means for their early detection. |
| Approach: | They propose to use a hierarchical attention network to predict if a user suffers from one of nine disorders to adapt a deep neural model to the task. |
| Outcome: | The proposed model outperforms previous benchmarks for four out of nine disorders in a binary classification task on social media. |
Representing Social Media Users for Sarcasm Detection (D18-1)
Copied to clipboard
| Challenge: | Existing annotated corpus of Reddit comments is limited by available annotation methods. |
| Approach: | They propose a Bayesian approach that directly represents authors’ propensities to be sarcastic and a dense embedding approach that can learn interactions between the author and the text. |
| Outcome: | The proposed approach performs better in homogeneous contexts, whereas the dense embeddings prove valuable in more diverse contexts. |
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models. |
| Approach: | They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups. |
| Outcome: | The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech. |