Challenge: Disclosure is a social media activity that can be rewarding but also poses privacy risks.
Approach: They propose to detect and abstract online self-disclosures using a large corpus of 4.8K annotated disclosure spans and a language model to fine-tune for detection.
Outcome: The proposed model can detect and abstract self-disclosures with 80% accuracy, on-par with GPT-3.5.

Similar Papers

A Semantics-based Approach to Disclosure Classification in User-Generated Online Content (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing algorithms for self-disclosure identification and classification are challenging due to the relative anonymity of social networking sites and lack of non-verbal cues to signal thoughts or feelings.
Approach: They propose an approach to detect emotional and informational self-disclosure in natural language by using frame semantics to identify lexical units and their semantic roles.
Outcome: The proposed method improves on reddit data and provides insights into the drivers of disclosure behaviors.
Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Media (2025.findings-emnlp)

Copied to clipboard

Challenge: Self-disclosure can provide psychological comfort but can also pose privacy concerns . a lack of high-quality corpora, analysis, and methods for detection is limiting research .
Approach: They construct a high-quality text-image corpus on Chinese multimodal social media platforms . they analyze the distribution of self-disclosure types, modality preferences, user intent .
Outcome: The proposed corpus analyzes self-disclosure behaviors on Chinese social media platforms . it fine-tunes five multimodal large language models to enhance self-discovery detection .
Examining the Utility of Self-disclosure Types for Modeling Annotators of Social Norms (2026.findings-eacl)

Copied to clipboard

Challenge: Recent work has explored the use of personal information in the form of persona sentences to improve modeling of individual characteristics and prediction of annotator labels for subjective tasks.
Approach: They categorize self-disclosures and use them to build annotator models for predicting judgments of social norms by analyzing comments from original post.
Outcome: The proposed model improves the model and its ability to predict annotator labels.
Measuring the Language of Self-Disclosure across Corpora (2022.findings-acl)

Copied to clipboard

Challenge: Existing models that estimate self-disclosure from language are poorly generalized due to variations in corpora and labeling instructions.
Approach: They build single-task models on five self-disclosure corpora and use them to predict self-declaration across corpors.
Outcome: The proposed model predicts self-disclosure across corpora, but the results are poor for out-of-corpora models.
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are accessed via commercial APIs, but expose data to service providers.
Approach: They propose a framework where a local model uses natural language instructions to rewrite queries and paired them with synthetic privacy profiles to achieve better privacy preservation.
Outcome: The proposed model outperforms large-scale few-shot models in terms of privacy preservation and performance.
Large Language Models Can Be Contextual Privacy Protection Learners (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable linguistic comprehension and generation capability, but when applied to specialized industries, they face challenges such as hallucination, insufficient domain knowledge, and failing to incorporate the latest domain knowledge.
Approach: They propose a paradigm for fine-tuning LLMs that effectively injects domain-specific knowledge while safeguarding inference-time data privacy.
Outcome: The proposed model protects private data while enhancing the model's knowledge.
Identifying Medical Self-Disclosure in Online Communities (2021.naacl-main)

Copied to clipboard

Challenge: a new dataset of health-related posts from online social platforms is available for analysis . medical self-disclosure may be useful for early detection and treatment of medical issues .
Approach: They propose to analyze medical self-disclosure in online health conversations . they release a dataset of health-related posts from online social platforms with high inter-annotator agreement .
Outcome: The proposed model achieves an accuracy of 81.02% and sets a strong performance benchmark.
Robust Utility-Preserving Text Anonymization Based on Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing techniques face challenges of re-identification ability of large language models . anonymizing text that contains sensitive information is crucial for a wide range of applications .
Approach: They propose a framework that integrates three key LLM components to perform anonymization.
Outcome: The proposed model outperforms baselines while maintaining greater data utility in downstream tasks.
How Private are Language Models in Abstractive Summarization? (2025.emnlp-main)

Copied to clipboard

Challenge: Effective protection of private information is essential for knowledge dissemination in sensitive domains such as medical and legal.
Approach: They perform a comprehensive study of privacy risks in LM-based summarization using closed- and four-weight models of different sizes and families.
Outcome: The proposed models show that they leak personally identifiable information in their summaries, compared to human-generated summary summators, which show significantly higher privacy protection levels.
Combating Security and Privacy Issues in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide a summary of risks and vulnerabilities in large language models . a number of studies have focused on security, privacy and copyright aspects of LLMs .
Approach: This tutorial seeks to provide a systematic summary of risks and vulnerabilities in large language models . authors will discuss security, privacy and copyright aspects of LLMs .
Outcome: This tutorial aims to provide a systematic summary of risks and vulnerabilities in large language models . it will also outline emerging challenges in security, privacy and reliability of LLMs .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations