The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)
Copied to clipboard
Salvatore Giorgi, Daniel Preoţiuc-Pietro, Anneke Buffone, Daniel Rieman, Lyle Ungar, H. Andrew Schwartz
| Challenge: | Social media data is often aggregated without regard to users in the Twitter populations of each community. |
| Approach: | They propose to use Twitter language to build community-level models using Twitter language aggregated by users. |
| Outcome: | The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets. |
Similar Papers
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)
Copied to clipboard
| Challenge: | Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias. |
| Approach: | They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC. |
| Outcome: | The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community. |
Residualized Factor Adaptation for Community Social Media Prediction Tasks (D18-1)
Copied to clipboard
| Challenge: | Existing approaches to social media language capture only socio-demographic contexts, such as age, education rates, race, and gender. |
| Approach: | They propose a method which integrates community attributes and adapts linguistic features to community attributes. |
| Outcome: | The proposed model integrates community attributes and adapts linguistic features to community attributes. |
Leakage-Aware User-Level ADHD Signal Classification from Social Media: When Graph Aggregation Helps, and When It Does Not (2026.acl-srw)
Copied to clipboard
| Challenge: | Social media data are longitudinal, usercentered, rich in spontaneous language use. |
| Approach: | They propose a leakage-aware evaluation framework organized around two controlled axes: evidence budget and leakage control. |
| Outcome: | The proposed framework compares graph aggregation with other models using psycholinguistic features and semantic tweet embeddings. |
TSix: A Human-involved-creation Dataset for Tweet Summarization (L18-1)
Copied to clipboard
| Challenge: | a new dataset for tweet summarization is available for free. |
| Approach: | They propose a dataset for tweet summarization that uses human annotations to evaluate extractive summarizing methods. |
| Outcome: | The proposed dataset includes six events collected from Twitter . human-annotated gold-standard references facilitate evaluation, the study shows . |
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models. |
| Approach: | They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups. |
| Outcome: | The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech. |
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Experimental results show improvements on Reddit and Twitter data . |
| Approach: | They propose to take advantage of Large Language Models (LLMs) to better identify user communities. |
| Outcome: | The proposed model improves on Reddit and Twitter data and tasks of community detection, bot detection, and news media profiling. |
Synthetic Data for English Lexical Normalization: How Close Can We Get to Manually Annotated Data? (2020.lrec-1)
Copied to clipboard
| Challenge: | Social media data is a valuable data resource for natural language processing tasks. |
| Approach: | They propose to adapt input text to a more standard form, a task also referred to as normalization. |
| Outcome: | The proposed system scores 94.29 accuracy on the test data compared to 95.22 when trained on human-annotated data. |
Twitter-Demographer: A Flow-based Tool to Enrich Twitter Data (2022.emnlp-demos)
Copied to clipboard
| Challenge: | 199 million people communicate on twitter daily, making it essential to study policy and decision-making. |
| Approach: | They propose a flow-based tool to augment Twitter data with additional information about tweets and users. |
| Outcome: | The proposed tool is designed to enhance Twitter data with additional information about tweets and users. |
TWEETSUM: Event oriented Social Summarization Dataset (2020.coling-main)
Copied to clipboard
| Challenge: | Developing social summarization systems is becoming more and more critical . but, the publicly available and high-quality large scale social summaries are rare . |
| Approach: | They propose to build a social summarization dataset using twitter's hot events . they collect user relations, hashtags and user profiles to evaluate their summarizing methods . |
| Outcome: | The proposed dataset is based on a dataset from twitter with 12 real world hot events with 44,034 tweets and 11,240 users. |
Label Embedding using Hierarchical Structure of Labels for Twitter Classification (D19-1)
Copied to clipboard
| Challenge: | Twitter is used for disaster monitoring and news material gathering . we propose a method that can consider the hierarchical structure of labels and labels themselves . |
| Approach: | They propose a method that can consider the hierarchical structure of labels and label texts themselves. |
| Outcome: | The proposed method outperforms the methods of the conference participants over the text REtrieval Conference (TREC) 2018 Incident Streams (IS) dataset. |