| Challenge: | Existing computational studies on stress only focus on domains such as speech or Twitter . a corpus of social media text is used to identify stress . |
| Approach: | They propose a text corpus of lengthy social media data for detecting stress . they use 190K posts from five different categories of Reddit communities . |
| Outcome: | The proposed corpus of social media data can be used to identify stress . it includes 190K posts from five different categories of Reddit communities . |
Similar Papers
MentalHelp: A Multi-Task Dataset for Mental Health in Social Media (2024.lrec-main)
Copied to clipboard
Nishat Raihan, Sadiya Sayara Chowdhury Puspo, Shafkat Farabi, Ana-Maria Bucur, Tharindu Ranasinghe, Marcos Zampieri
| Challenge: | Annotating social media data for mental health disorders is expensive and time-consuming, limiting their size and scope. |
| Approach: | They present a large-scale semi-supervised mental disorder detection dataset containing 14 million instances from Reddit and an ensemble of three separate models. |
| Outcome: | The proposed dataset contains 14 million instances of mental disorders . it was collected from reddit and labeled in a semi-supervised way . |
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models. |
| Approach: | They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups. |
| Outcome: | The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech. |
Weakly-Supervised Methods for Suicide Risk Assessment: Role of Related Domains (2021.acl-short)
Copied to clipboard
| Challenge: | Among social media platforms, Reddit has emerged as the most promising one due to its anonymity and its focus on topic-based communities (subreddits) . a challenge for previous work on suicide risk assessment has been the small amount of labeled data. |
| Approach: | They propose to use social media to collect user data from r/SuicideWatch subreddit and annotate it with user-level suicide risk: no-risk, low-risk and high-risk. |
| Outcome: | The proposed model improves by using pseudo-labeling based on related issues around mental health (e.g., anxiety, depression) |
RedDust: a Large Reusable Dataset of Reddit User Traits (2020.lrec-1)
Copied to clipboard
| Challenge: | Social media is a rich source of assertions about personal traits, but identifying personal traits from implicit assertions is difficult because of the users’ highly varied vocabulary and expressions. |
| Approach: | They propose to build a large-scale annotated resource for user profiling for over 300k Reddit users across five attributes: profession, hobby, family status, age, and gender. |
| Outcome: | The proposed resource is the first annotated language resource about Reddit users at large scale. |
MentSum: A Resource for Exploring Summarization of Mental Health Online Posts (2022.lrec-1)
Copied to clipboard
| Challenge: | Mental health remains a significant challenge of public health worldwide . many use online platforms to share their mental health conditions and seek help . |
| Approach: | They analyze a dataset of over 24k user posts from Reddit and 43 mental health subreddits to generate a short summarization. |
| Outcome: | The proposed dataset compared over 24k user posts and 43 mental health subreddits . it shows that the summarization of these posts is faster and more accurate than previous studies. |
Introducing CAD: the Contextual Abuse Dataset (2021.naacl-main)
Copied to clipboard
| Challenge: | Detecting and classifying online abuse is a complex and nuanced task, despite many advances in the power and availability of computational tools. |
| Approach: | They propose to annotate a reddit conversation thread with six distinct primary and secondary categories and an expert-driven group-adjudication process for high quality annotations. |
| Outcome: | The proposed dataset contains six distinct primary and secondary categories and uses an expert-driven group-adjudication process for high quality annotations. |
An Annotated Dataset for Explainable Interpersonal Risk Factors of Mental Disturbance in Social Media Posts (2023.findings-acl)
Copied to clipboard
| Challenge: | 1.6 million people in England are on waiting lists for mental health care . 8 million people are not considered sick enough to qualify for help . |
| Approach: | They construct and release an annotated dataset with human-labelled explanations and classification of IRF affecting mental disturbance on social media: TBe, PBu, and PBe. |
| Outcome: | The proposed model can be used to identify TBe and PBU in emotional spectrum of user’s historical social media profile. |
RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media (2023.findings-eacl)
Copied to clipboard
| Challenge: | Social media platforms such as Reddit are vulnerable to misinformation and disinformation. |
| Approach: | They propose a method to automatically derive (noisy) supervision for retrieval of trustworthy evidence relevant to a given claim made on social media. |
| Outcome: | The proposed method outperforms baseline models in the retrieval task performed by medical doctors. |
Do Models of Mental Health Based on Social Media Data Generalize? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing literature on the validity of proxy-based methods for annotating mental health status in social media has raised new concerns regarding their use in clinical applications. |
| Approach: | They explore the generalization ability of machine learning classifiers trained to detect depression in individuals across multiple social media platforms. |
| Outcome: | The proposed methods show that they can be used to train and analyze large datasets and that they are robust to large dataset sizes. |
A System for Dynamically Tracking Content Moderation on Reddit (2026.acl-demo)
Copied to clipboard
| Challenge: | Recent work in social media platforms delegate content moderation decisions to users and communities. |
| Approach: | They propose a software system for the dynamic monitoring of Reddit posts, communities, and moderation actions to enable scalable and reproducible research on decentralized platform governance and content moderation. |
| Outcome: | The proposed system is the only available solution for general-purpose, real-time, policy-compliant longitudinal data collection on Reddit. |