Challenge: Social media data is often aggregated without regard to users in the Twitter populations of each community.
Approach: They propose to use Twitter language to build community-level models using Twitter language aggregated by users.
Outcome: The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets.

Similar Papers

User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)

Copied to clipboard

Challenge: Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias.
Approach: They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC.
Outcome: The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community.
Residualized Factor Adaptation for Community Social Media Prediction Tasks (D18-1)

Copied to clipboard

Challenge: Existing approaches to social media language capture only socio-demographic contexts, such as age, education rates, race, and gender.
Approach: They propose a method which integrates community attributes and adapts linguistic features to community attributes.
Outcome: The proposed model integrates community attributes and adapts linguistic features to community attributes.
Leakage-Aware User-Level ADHD Signal Classification from Social Media: When Graph Aggregation Helps, and When It Does Not (2026.acl-srw)

Copied to clipboard

Challenge: Social media data are longitudinal, usercentered, rich in spontaneous language use.
Approach: They propose a leakage-aware evaluation framework organized around two controlled axes: evidence budget and leakage control.
Outcome: The proposed framework compares graph aggregation with other models using psycholinguistic features and semantic tweet embeddings.
TSix: A Human-involved-creation Dataset for Tweet Summarization (L18-1)

Copied to clipboard

Challenge: a new dataset for tweet summarization is available for free.
Approach: They propose a dataset for tweet summarization that uses human annotations to evaluate extractive summarizing methods.
Outcome: The proposed dataset includes six events collected from Twitter . human-annotated gold-standard references facilitate evaluation, the study shows .
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models.
Approach: They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups.
Outcome: The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech.
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media (2024.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show improvements on Reddit and Twitter data .
Approach: They propose to take advantage of Large Language Models (LLMs) to better identify user communities.
Outcome: The proposed model improves on Reddit and Twitter data and tasks of community detection, bot detection, and news media profiling.
Synthetic Data for English Lexical Normalization: How Close Can We Get to Manually Annotated Data? (2020.lrec-1)

Copied to clipboard

Challenge: Social media data is a valuable data resource for natural language processing tasks.
Approach: They propose to adapt input text to a more standard form, a task also referred to as normalization.
Outcome: The proposed system scores 94.29 accuracy on the test data compared to 95.22 when trained on human-annotated data.
Twitter-Demographer: A Flow-based Tool to Enrich Twitter Data (2022.emnlp-demos)

Copied to clipboard

Challenge: 199 million people communicate on twitter daily, making it essential to study policy and decision-making.
Approach: They propose a flow-based tool to augment Twitter data with additional information about tweets and users.
Outcome: The proposed tool is designed to enhance Twitter data with additional information about tweets and users.
TWEETSUM: Event oriented Social Summarization Dataset (2020.coling-main)

Copied to clipboard

Challenge: Developing social summarization systems is becoming more and more critical . but, the publicly available and high-quality large scale social summaries are rare .
Approach: They propose to build a social summarization dataset using twitter's hot events . they collect user relations, hashtags and user profiles to evaluate their summarizing methods .
Outcome: The proposed dataset is based on a dataset from twitter with 12 real world hot events with 44,034 tweets and 11,240 users.
Label Embedding using Hierarchical Structure of Labels for Twitter Classification (D19-1)

Copied to clipboard

Challenge: Twitter is used for disaster monitoring and news material gathering . we propose a method that can consider the hierarchical structure of labels and labels themselves .
Approach: They propose a method that can consider the hierarchical structure of labels and label texts themselves.
Outcome: The proposed method outperforms the methods of the conference participants over the text REtrieval Conference (TREC) 2018 Incident Streams (IS) dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations