Challenge: Existing methods to train federated learning (FL) for natural language processing require sensitive data to leave local devices.
Approach: They propose a fedrated model decomposition method that protects the privacy of vocabularies . they propose an adaptive updating technique to improve the performance of local models .
Outcome: The proposed method protects the privacy of vocabularies in federated learning tasks . it maintains competitive performance and provides better privacy-preserving capacity compared to status quo methods.

Similar Papers

Safely Learning with Private Data: A Federated Learning Framework for Large Language Model (2024.emnlp-main)

Copied to clipboard

Challenge: Existing large language models (LLMs) use large amounts of public data and massive parameters, but private data is often stored in isolated data silos.
Approach: They propose a Federated Learning framework for large language models which offloads most training parameters to the server while training embedding and output layers locally.
Outcome: The proposed framework achieves comparable metrics to centralized chatGLM model on NLU and generation tasks.
Empirical Studies of Institutional Federated Learning For Natural Language Processing (2020.findings-emnlp)

Copied to clipboard

Challenge: federated learning is a promising ideology to unite isolated datasets for machine learning problems.
Approach: They propose to use federated natural language processing networks to train a popular NLP model with applications in sentence intent classification.
Outcome: The proposed model is sensitive to imbalanced data load and tested against a federated model under imbalanced datasets.
FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks (2022.findings-naacl)

Copied to clipboard

Challenge: Increasing concerns and regulations about data privacy necessitate the study of privacy-preserving, decentralized learning methods for natural language processing tasks.
Approach: They propose a framework for evaluating federated learning methods on four different tasks . they propose federation between Transformer-based language models and FL methods .
Outcome: The proposed framework compares FL methods on four different tasks under non-IID partitioning strategies.
Efficient Federated Learning on Knowledge Graphs via Privacy-preserving Relation Embedding Aggregation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing frameworks that share entity embeddings of knowledge graphs (KGs) would incur a severe privacy leakage.
Approach: They propose a new attack method that aims to recover the original embedding information based on the known entity embeddables of FedE.
Outcome: The proposed framework can be used to infer whether a specific relation exists in a private client.
Privacy-Preserving Natural Language Processing (2023.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial will help the NLP community to get familiar with current research in privacy-preserving methods.
Approach: This tutorial will help the NLP community to get familiar with current research in privacy-preserving methods.
Outcome: The tutorial will cover membership inference, differential privacy, homomorphic encryption, or federated learning, all with typical use-cases and potential pitfalls.
Practical Takes on Federated Learning with Pretrained Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: federated learning with pretrained language models for language tasks entails data privacy constraints when learning from diverse data domains.
Approach: They propose to use pretrained language models to learn from diverse data domains . they elaborate hypotheses over the components in federated NLP architectures based on three tasks .
Outcome: The proposed model can generalize by adapting to the different domains.
Revisiting Data Reconstruction Attacks on Real-world Dataset for Federated Natural Language Understanding (2024.lrec-main)

Copied to clipboard

Challenge: Existing DRA methods fail to accurately recover the original text of real-world privacy data.
Approach: They propose to use a real-world privacy dataset to examine the performance of federated learning (FL) methods.
Outcome: The proposed method improves on a real-world privacy dataset and shows that the tokens within a recovery sentence are disordered and intertwined with tokens from other sentences in the same training batch.
Fair NLP Models with Differentially Private Text Encoders (2022.findings-emnlp)

Copied to clipboard

Challenge: Encoded text representations often capture sensitive attributes about individuals, raising privacy concerns and making models unfair to certain groups.
Approach: They propose an approach that combines privacy and adversarial training to learn private representations which induces fairer models.
Outcome: The proposed approach improves on four NLP datasets and shows that privacy and fairness can positively reinforce each other.
Privacy-preserving Neural Representations of Text (D18-1)

Copied to clipboard

Challenge: a specific type of attack is used to characterize the privacy of neural representations for NLP tasks, in the context of privacy protection.
Approach: They propose several defense methods based on modified training objectives and characterize the tradeoff between privacy and the utility of neural representations.
Outcome: The proposed defenses improve the privacy of neural representations and characterize the tradeoff between privacy and utility of representations.
Gradient Inversion Attack in Federated Learning: Exposing Text Data through Discrete Optimization (2025.coling-main)

Copied to clipboard

Challenge: federated learning could overcome the bottleneck of public text data in large language models . a novel attack method is proposed to fully expose text data from gradients .
Approach: They propose a method to fully expose text data from gradients by using a network of clients and a server.
Outcome: The proposed method shows it is possible to Fully Expose Text data from gradients.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations