Challenge: Pretrained language models (PLMs) are used for personalized federated learning . communication costs are high with large PLMs, and local training is expensive .
Approach: They propose a framework for federated learning with pretrained language models . they propose 'discrete local search' and compression mechanism for local training .
Outcome: The proposed framework achieves superior performance compared with baselines.

Similar Papers

FedPETuning: When Federated Learning Meets the Parameter-Efficient Tuning Methods of Pre-trained Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Existing research on federated learning (FL) for pre-trained language models (PLMs) with increasing concerns about data privacy, enterprises or institutions are not allowed to collect data from end devices or local clients to a centralized server for fine-tuning PLMs.
Approach: They investigate the parameter-efficient tuning of pre-trained language models (PLMs) and develop a federated benchmark for four representative PETuning methods .
Outcome: The proposed method can defend against privacy attacks and maintain acceptable performance with reducing heavy resource consumption.
Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization (2023.emnlp-main)

Copied to clipboard

Challenge: Prompt tuning of Large Language Models (LLMs) can incur performance degradation or low training efficiency.
Approach: They propose a prompt tuning approach with Adaptive Optimization to enable efficient FL of LLMs.
Outcome: The proposed approach improves performance and efficiency simultaneously and addresses client drift problems on both the device and server sides.
Reliable Gradient-free and Likelihood-free Prompt Tuning (2023.findings-eacl)

Copied to clipboard

Challenge: Large pre-trained language models are often offered as black-box APIs due to privacy or commercial constraints.
Approach: They propose to tune the soft prompts without requiring gradient computation and extend the model to include a distribution over prompts.
Outcome: The proposed methods are competitive with gradient-based approaches with full access to the PLM.
FedPerC: Federated Learning for Language Generation with Personal and Context Preference Embeddings (2023.findings-eacl)

Copied to clipboard

Challenge: federated learning is a decentralized learning paradigm that assumes no access to a large labeled dataset and instead leverages averaged parameter updates across all users of the system.
Approach: They propose a method to personalize federated learning with personal embeddings and shared context embeddables.
Outcome: The proposed approach achieves 50% improvement in test-time perplexity using 0.001% of the memory required by baseline approaches and greater sample- and compute-efficiency.
Global Adaptive Momentum Meets Local Personalized Perturbation: Efficient Federated LLM Fine-Tuning with Zeroth-Order Gradients (2026.acl-long)

Copied to clipboard

Challenge: federated fine-tuning of large language models provides privacy-preserving approach to deploying pervasive generative AI services.
Approach: They propose a federated framework for fine-tuning large language models . they propose unified optimization and local personalized perturbation for ZO gradients .
Outcome: The proposed framework outperforms existing methods for integrating ZO gradients in federated learning over diverse heterogeneous data settings.
Practical Takes on Federated Learning with Pretrained Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: federated learning with pretrained language models for language tasks entails data privacy constraints when learning from diverse data domains.
Approach: They propose to use pretrained language models to learn from diverse data domains . they elaborate hypotheses over the components in federated NLP architectures based on three tasks .
Outcome: The proposed model can generalize by adapting to the different domains.
Gradient Inversion Attack in Federated Learning: Exposing Text Data through Discrete Optimization (2025.coling-main)

Copied to clipboard

Challenge: federated learning could overcome the bottleneck of public text data in large language models . a novel attack method is proposed to fully expose text data from gradients .
Approach: They propose a method to fully expose text data from gradients by using a network of clients and a server.
Outcome: The proposed method shows it is possible to Fully Expose Text data from gradients.
Client-Customized Adaptation for Parameter-Efficient Federated Learning (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models have a large memory footprint and are difficult to use in federated learning (FL)
Approach: They propose a hypernetwork-based FL framework that generates client-specific adapters by conditioning the client information.
Outcome: The proposed framework maximizes the utility of shared model parameters while minimizing divergence caused by client heterogeneity.
Communication-Efficient and Tensorized Federated Fine-Tuning of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in translation and summarization due to the capabilities of transformer architectures.
Approach: They propose to integrate tensorized adapters into model encoder/decoder blocks to improve model adaptability against data heterogeneity.
Outcome: Experiments on large-scale cross-device FL and large-silo FL show that the proposed methods perform on par or even better than existing federated PEFT approaches while reducing communication cost.
BBTv2: Towards a Gradient-Free Future with Large Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Recent work on parameter-efficient tuning (PET) only tunes a small portion of parameters while keeping most of the parameters of the LLM unchanged.
Approach: They propose an improved version of Black-Box Tuning to tune PTMs through gradient descent . they prepend continuous prompts to every layer of the PTM and propose a divide-and-conquer gradient-free algorithm to optimize the prompts alternately.
Outcome: The proposed method achieves comparable performance to full model tuning and state-of-the-art parameter-efficient methods under few-shot settings while maintaining much fewer tunable parameters.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations