Challenge: TextHide is a proposed privacy-enhancing technology to protect the training data from privacy attacks.
Approach: They propose to encode training data via instance encoding in natural language domain without theoretic privacy guarantee.
Outcome: The proposed scheme can defend against privacy attacks while ensuring learning utility (as a trade-off).

Similar Papers

Reconstruction Attack on Instance Encoding for Language Understanding (2021.emnlp-main)

Copied to clipboard

Challenge: Existing private learning schemes which protect data privacy can be used to train models using instance encoding.
Approach: They propose to recover the private training data and use it to break a private learning scheme TextHide.
Outcome: The proposed attack would advance privacy-preserving machine learning in the context of natural language processing.
ADePT: Auto-encoder based Differentially Private Text Transformation (2021.eacl-main)

Copied to clipboard

Challenge: Differential privacy is an important privacy concern when building statistical models on data containing sensitive information.
Approach: They propose a utility-preserving differentially private text transformation algorithm using auto-encoders that can be used to transform text to offer robustness against attacks and produce transformations with high semantic quality.
Outcome: The proposed model performs better against membership inference attacks while offering lower to no degradation in the utility of the underlying transformation process compared to baselines.
TextHide: Tackling Data Privacy in Language Understanding Tasks (2020.findings-emnlp)

Copied to clipboard

Challenge: Unsolved privacy challenges in distributed or federated learning are a challenge for many domains including Natural Language Processing.
Approach: They propose a federated learning framework that adds an encryption step to prevent an eavesdropping attacker from recovering private text data.
Outcome: The proposed model can effectively defend against attacks on shared gradients or representations and the averaged accuracy reduction is only 1.9%.
How reparametrization trick broke differentially-private text representation learning (2022.acl-short)

Copied to clipboard

Challenge: Differential privacy (DP) is a formal mathematical treatment of privacy protection . it guarantees how much privacy can be lost in the worst case . adapting DP mechanisms to NLP properly is largely non-trivial task .
Approach: They propose to use differential privacy to learn text representations using DPText to quantify and guarantee how much privacy can be lost in the worst case.
Outcome: The proposed methods are falsely claimed to be differentially private and violate privacy loss guarantees.
Differentially Private Representation for NLP: Formal Guarantee and An Empirical Study on Privacy and Fairness (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to learn text representations can encode private information of the input, thus can be exploited to recover such information with reasonable accuracy.
Approach: They propose a novel approach to preserve privacy of the extracted representation from text by combining differential privacy with dropout.
Outcome: The proposed approach preserves privacy of the extracted representation from text while masking words via dropout can enhance privacy.
Differentially Private Language Models for Secure Data Sharing (2022.emnlp-main)

Copied to clipboard

Challenge: a variety of deanonymization attacks allow the re-identification of individuals from tabular data.
Approach: They propose to train a language model in a differentially private manner and sample data from it . they find that the model generates fluent textual datasets with privacy guarantees .
Outcome: The proposed methods outperform direct classifiers with DP-SGD in the real-world.
When differential privacy meets NLP: The devil is in the detail (2021.emnlp-main)

Copied to clipboard

Challenge: Differential privacy provides a formal approach to privacy of individuals.
Approach: They propose to use ADePT to provide differentially private auto-encoders for text rewriting to provide tight privacy guarantees for users' original utterances.
Outcome: The proposed algorithm is not differentially private, thus rendering the experimental results unsubstantiated.
Synthetic Text Generation with Differential Privacy: A Simple and Practical Recipe (2023.acl-long)

Copied to clipboard

Challenge: Privacy concerns have increased in data-driven products due to the tendency of machine learning models to memorize sensitive training data.
Approach: They propose a method for generating useful synthetic text with a formal privacy guarantee by fine-tuning a pretrained generative language model with DP.
Outcome: The proposed method produces synthetic text competitive in terms of utility with its non-private counterpart, while providing strong protection against potential privacy leakages.
Privacy-preserving Neural Representations of Text (D18-1)

Copied to clipboard

Challenge: a specific type of attack is used to characterize the privacy of neural representations for NLP tasks, in the context of privacy protection.
Approach: They propose several defense methods based on modified training objectives and characterize the tradeoff between privacy and the utility of neural representations.
Outcome: The proposed defenses improve the privacy of neural representations and characterize the tradeoff between privacy and utility of representations.
The Limits of Word Level Differential Privacy (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to anonymize textual data have several shortcomings . authors show that they can overcome these weaknesses and offer a formal privacy guarantee .
Approach: They propose a method that circumvents most of the identified weaknesses and offers a formal privacy guarantee.
Outcome: The proposed method outperforms the proposed methods in thourough experimentation and shows superior performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations