Learning Invariant Representations of Social Media Users (D19-1)

Copied to clipboard

Challenge: Existing methods for learning to compare social media users fail to generalize to new users or even to previously known users.
Approach: They propose a procedure to learn a mapping from short episodes of user activity to a vector space in which the distance between points captures the similarity of the corresponding users’ invariant features.
Outcome: The proposed procedure can be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space.

Similar Papers

A Deep Metric Learning Approach to Account Linking (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to identify abusive content may fail to adapt to new trends, and individual posts may fail .
Approach: They propose a method that embeds variable-sized samples of user activity into a vector space, where samples by the same author map to nearby points.
Outcome: The proposed model outperforms several competitive baselines under a new evaluation framework modeled after established benchmarks in other domains.
You Shall Know a User by the Company It Keeps: Dynamic Representations for Social Media Users in NLP (D19-1)

Copied to clipboard

Challenge: Current approaches to social media modelling ignore the fact that an individual may be part of several communities which are not equally relevant in all communicative situations.
Approach: They propose a model that captures the sociological phenomenon of homophily and combines it with linguistic information to make a prediction.
Outcome: The proposed model significantly outperforms existing models on three different tasks and is compared with other models.
Simple Attention-Based Representation Learning for Ranking Short Social Media Posts (N19-1)

Copied to clipboard

Challenge: Existing approaches to ranking short social media posts are complex and require different components to capture a multitude of relevance signals.
Approach: They propose a word-level Siamese architecture with attention-based mechanisms for capturing semantic "soft" matches between query and post tokens.
Outcome: The proposed model is faster and simpler than existing models and more efficient than existing approaches.
PASUM: A Pre-training Architecture for Social Media User Modeling Based on Text Graph (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have incorporated different digital traces to better learn the representations of social media users, limited by overloaded text information and hard-to-collect social network information.
Approach: They propose a Pre-training Architecture for Social Media User Modeling based on Text Graph and combine microblogs to represent social media users based upon the text graph model.
Outcome: The proposed framework can represent users based on text even without social network information on microblogs.
Creation and evaluation of timelines for longitudinal user posts (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods for segmenting user posts into timelines improve quality and cost of manual annotation.
Approach: They propose a set of methods for segmenting longitudinal user posts into timelines likely to contain interesting moments of change in a user’s behaviour based on their online posting activity.
Outcome: The proposed framework is able to evaluate two different social media datasets and compares with existing models.
Cross-media User Profiling with Joint Textual and Social User Embedding (C18-1)

Copied to clipboard

Challenge: Empirical studies demonstrate the effectiveness of the proposed approach to cross-media user profiling tasks.
Approach: They propose a uniform user embedding learning approach to address cross-media user profiling by bridging the knowledge between the source and target media.
Outcome: Empirical results show that the proposed approach performs well on two cross-media user profiling tasks.
Extracting Age-Related Stereotypes from Social Media Texts (2022.lrec-1)

Copied to clipboard

Challenge: a method for extracting age-related stereotypes from Twitter data is under-studied in NLP . stereotyping on the basis of protected characteristics has been understudied .
Approach: They propose a method for extracting age-related stereotypes from Twitter data . they generate a corpus of 300,000 over-generalizations about four contemporary generations .
Outcome: The method uncovers common stereotypes as reported in media and psychological literature . it also finds that stereotypes for different generations vary across topics .
Adapting Deep Learning Methods for Mental Health Prediction on Social Media (D19-55)

Copied to clipboard

Challenge: a quarter of the population in Europe suffers from an episode of a mental disorder in their life, according to the World Health Organization . text analysis of rich resources like social media can contribute to deeper understanding of mental health and provide means for their early detection.
Approach: They propose to use a hierarchical attention network to predict if a user suffers from one of nine disorders to adapt a deep neural model to the task.
Outcome: The proposed model outperforms previous benchmarks for four out of nine disorders in a binary classification task on social media.
Representing Social Media Users for Sarcasm Detection (D18-1)

Copied to clipboard

Challenge: Existing annotated corpus of Reddit comments is limited by available annotation methods.
Approach: They propose a Bayesian approach that directly represents authors’ propensities to be sarcastic and a dense embedding approach that can learn interactions between the author and the text.
Outcome: The proposed approach performs better in homogeneous contexts, whereas the dense embeddings prove valuable in more diverse contexts.
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models.
Approach: They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups.
Outcome: The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations