MentSum: A Resource for Exploring Summarization of Mental Health Online Posts (2022.lrec-1)
Copied to clipboard
| Challenge: | Mental health remains a significant challenge of public health worldwide . many use online platforms to share their mental health conditions and seek help . |
| Approach: | They analyze a dataset of over 24k user posts from Reddit and 43 mental health subreddits to generate a short summarization. |
| Outcome: | The proposed dataset compared over 24k user posts and 43 mental health subreddits . it shows that the summarization of these posts is faster and more accurate than previous studies. |
Similar Papers
Leveraging Language Models for Summarizing Mental State Examinations: A Comprehensive Evaluation and Dataset Release (2025.coling-main)
Copied to clipboard
| Challenge: | Mental health disorders affect a significant portion of the global population . access to mental health support is limited in developing countries . |
| Approach: | They evaluated a 12-item descriptive MSE questionnaire and five well-known summarization models . they found that language models can generate coherent MSE summaries for doctors . |
| Outcome: | The proposed model can generate coherent summaries from MSEs in a conversational format. |
WikiSum: Coherent Summarization Dataset for Efficient Human-Evaluation (2021.acl-short)
Copied to clipboard
| Challenge: | Existing summarization datasets are limited in their ability to evaluate output . a human evaluation is necessary to understand and improve summarizing systems . |
| Approach: | They propose a dataset based on how-to articles and coherent paragraph summaries written in plain language. |
| Outcome: | The proposed dataset makes human evaluation easier and more effective . the authors compare the proposed dataset to existing ones on PubMed and the literature. |
SumPubMed: Summarization Dataset of PubMed Scientific Articles (2021.acl-srw)
Copied to clipboard
| Challenge: | Existing summarization models that can extract the top few lines of news articles fail to summarize long documents. |
| Approach: | They constructed a scientific summarization dataset from MEDLINE articles from the PubMed archive to address this problem. |
| Outcome: | The proposed model outperforms existing models on news article summarization datasets and shows that it is more efficient to extract the top few lines. |
Bringing Structure into Summaries: a Faceted Summarization Dataset for Long Scientific Documents (2021.acl-short)
Copied to clipboard
| Challenge: | Faceted summarization provides briefings of a document from different perspectives. |
| Approach: | They propose a faceted summarization benchmark built on Emerald journal articles . they propose faceted models that bring structure into faceted documents . |
| Outcome: | The proposed benchmark is based on Emerald journal articles and covers a diverse range of domains. |
How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure Evaluation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods to evaluate summary coherence are often evaluated using disparate datasets and metrics. |
| Approach: | They propose to use automatic evaluation to evaluate coherence of summaries by selecting high-scoring candidates. |
| Outcome: | The proposed methods show that they can perform better on an even playing field. |
SummEval: Re-evaluating Summarization Evaluation (2021.tacl-1)
Copied to clipboard
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, Dragomir Radev
| Challenge: | a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments . |
| Approach: | They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization . |
| Outcome: | The proposed evaluation metrics are inconsistent with existing evaluation protocols. |
MentalHelp: A Multi-Task Dataset for Mental Health in Social Media (2024.lrec-main)
Copied to clipboard
Nishat Raihan, Sadiya Sayara Chowdhury Puspo, Shafkat Farabi, Ana-Maria Bucur, Tharindu Ranasinghe, Marcos Zampieri
| Challenge: | Annotating social media data for mental health disorders is expensive and time-consuming, limiting their size and scope. |
| Approach: | They present a large-scale semi-supervised mental disorder detection dataset containing 14 million instances from Reddit and an ensemble of three separate models. |
| Outcome: | The proposed dataset contains 14 million instances of mental disorders . it was collected from reddit and labeled in a semi-supervised way . |
TWEETSUM: Event oriented Social Summarization Dataset (2020.coling-main)
Copied to clipboard
| Challenge: | Developing social summarization systems is becoming more and more critical . but, the publicly available and high-quality large scale social summaries are rare . |
| Approach: | They propose to build a social summarization dataset using twitter's hot events . they collect user relations, hashtags and user profiles to evaluate their summarizing methods . |
| Outcome: | The proposed dataset is based on a dataset from twitter with 12 real world hot events with 44,034 tweets and 11,240 users. |
Weakly-Supervised Methods for Suicide Risk Assessment: Role of Related Domains (2021.acl-short)
Copied to clipboard
| Challenge: | Among social media platforms, Reddit has emerged as the most promising one due to its anonymity and its focus on topic-based communities (subreddits) . a challenge for previous work on suicide risk assessment has been the small amount of labeled data. |
| Approach: | They propose to use social media to collect user data from r/SuicideWatch subreddit and annotate it with user-level suicide risk: no-risk, low-risk and high-risk. |
| Outcome: | The proposed model improves by using pseudo-labeling based on related issues around mental health (e.g., anxiety, depression) |
How well do you know your summarization datasets? (2021.findings-acl)
Copied to clipboard
| Challenge: | State-of-the-art summarization systems are trained on massive datasets scraped from the web. |
| Approach: | They manually analyse 600 samples from three popular summarization datasets . they use a six-class typology which captures different noise types and degrees of summarizing difficulty. |
| Outcome: | The proposed model performs better on large datasets than on the current models. |