Challenge: Sense of Community is a social motivation that is reflected in the social behavior of humans.
Approach: They compile a large collection of parallel community datasets comprising over 7 million posts and comments from Reddit and 200,000 posts and comment from Dread, a dark web discussion forum, covering similar topics.
Outcome: The results show that users on Reddit exhibit a stronger sense of community membership despite the dark web’s restricted accessibility.

Similar Papers

Dreaddit: A Reddit Dataset for Stress Analysis in Social Media (D19-62)

Copied to clipboard

Challenge: Existing computational studies on stress only focus on domains such as speech or Twitter . a corpus of social media text is used to identify stress .
Approach: They propose a text corpus of lengthy social media data for detecting stress . they use 190K posts from five different categories of Reddit communities .
Outcome: The proposed corpus of social media data can be used to identify stress . it includes 190K posts from five different categories of Reddit communities .
Investigating Online Community Engagement through Stancetaking (2023.findings-emnlp)

Copied to clipboard

Challenge: Large-scale computational work on stancetaking has explored community similarities in their preferences for stance markers without considering the stance-relevant properties of the contexts in which stance marker use is carried out.
Approach: They propose to use stance-relevant properties of Reddit communities to capture community identity patterns distinct from textual or marker similarity measures.
Outcome: The proposed representations capture community identity patterns distinct from textual or marker similarity measures and relate them to broader inter- and intra-community engagement patterns.
SYSML: StYlometry with Structure and Multitask Learning: Implications for Darknet Forum Migrant Analysis (2021.emnlp-main)

Copied to clipboard

Challenge: Crypto markets are forums where goods and services are exchanged between parties who use encryption to conceal their identities.
Approach: They propose a stylometry-based multitask learning approach for natural language and model interactions using graph embeddings.
Outcome: The proposed approach outperforms existing methods in four darknet forums with a lift of up to 2.5X on the mean retrieval rank and 2X on recall@10.
Investigating Human Values in Online Communities (2025.naacl-long)

Copied to clipboard

Challenge: Existing value frameworks struggle with sample sizes and rely on selfreported surveys to calculate values.
Approach: They propose a method to computationally analyse values on Reddit using in-domain and out-of-domain human annotations to train a value relevance and a polarity classifier.
Outcome: The proposed method can be used to analyse values on reddit using human annotations and human annotation.
MentalHelp: A Multi-Task Dataset for Mental Health in Social Media (2024.lrec-main)

Copied to clipboard

Challenge: Annotating social media data for mental health disorders is expensive and time-consuming, limiting their size and scope.
Approach: They present a large-scale semi-supervised mental disorder detection dataset containing 14 million instances from Reddit and an ensemble of three separate models.
Outcome: The proposed dataset contains 14 million instances of mental disorders . it was collected from reddit and labeled in a semi-supervised way .
Analyzing Hate Speech Amplification on Fringe Platforms (2026.acl-srw)

Copied to clipboard

Challenge: a new study examines the factors that determine how hate speech amplifies on fringe platforms . fringe platforms like Gab harbor high volumes of hate speech due to minimal moderation and insular communities .
Approach: They used a dataset of 5K+ threads and 50K+ responses from four fringe platforms . they found that thread structure and disagreements in early response windows can give up to 74% lift in RMSE .
Outcome: The proposed model estimates how several features influence hate speech amplification on fringe platforms.
SOBR: A Corpus for Stylometry, Obfuscation, and Bias on Reddit (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora are limited in scope and can be used to collect data on author attributes.
Approach: They propose to use subreddits, flairs, and self-reports as distant labels for author attributes (age, gender, nationality, personality, and political leaning) .
Outcome: The proposed method could be used to infer author attributes from public posts despite their discreetness and anonymity .
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)

Copied to clipboard

Challenge: Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools.
Approach: They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations.
Outcome: The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations.
A System for Dynamically Tracking Content Moderation on Reddit (2026.acl-demo)

Copied to clipboard

Challenge: Recent work in social media platforms delegate content moderation decisions to users and communities.
Approach: They propose a software system for the dynamic monitoring of Reddit posts, communities, and moderation actions to enable scalable and reproducible research on decentralized platform governance and content moderation.
Outcome: The proposed system is the only available solution for general-purpose, real-time, policy-compliant longitudinal data collection on Reddit.
ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions (2024.findings-emnlp)

Copied to clipboard

Challenge: VALUESCOPE is a framework that quantifies social norms and values within online communities.
Approach: They propose a framework that uses language models to quantify social norms and values within online communities.
Outcome: The proposed framework delineates differences in social norms and tracks evolution of norms in online communities and influence of significant external events like the U.S. presidential elections and the emergence of new sub-communities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations