Papers by Christophe Dupuy

6 papers
Coordinated Replay Sample Selection for Continual Federated Learning (2023.emnlp-industry)

Copied to clipboard

Challenge: Continual Federated Learning (CFL) combines decentralized learning with continuous learning . ubiquity of personal devices with a network connection offers rich source of data for learning problems .
Approach: They propose to combine decentralized learning with a continuous learning approach . they propose to coordinate gradient-based replay sample selection across clients .
Outcome: The proposed method shows gains early in the low replay size regime, when the budget for storing past data is small.
FLIRT: Feedback Loop In-context Red Teaming (2024.emnlp-main)

Copied to clipboard

Challenge: Recent work has evaluated the vulnerabilities of large generative models, such as DALL-E, ChatGPT, and GPT-4.
Approach: They propose an automatic red teaming framework that evaluates a given black-box model and exposes its vulnerabilities against unsafe and inappropriate content generation.
Outcome: The proposed framework evaluates a given black-box model and exposes its vulnerabilities against unsafe and inappropriate content generation.
FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks (2022.findings-naacl)

Copied to clipboard

Challenge: Increasing concerns and regulations about data privacy necessitate the study of privacy-preserving, decentralized learning methods for natural language processing tasks.
Approach: They propose a framework for evaluating federated learning methods on four different tasks . they propose federation between Transformer-based language models and FL methods .
Outcome: The proposed framework compares FL methods on four different tasks under non-IID partitioning strategies.
Canary Extraction in Natural Language Understanding Models (2022.acl-short)

Copied to clipboard

Challenge: Recent literature has focused on Model Inversion Attacks (ModIvA) that can extract training data from model parameters.
Approach: They propose an attack that extracts canaries from NLU training data and reconstructs them using non-sensitive tokens.
Outcome: The proposed attack can reconstruct a four digit code in the training dataset with a probability of 0.5 in its best configuration.
ADePT: Auto-encoder based Differentially Private Text Transformation (2021.eacl-main)

Copied to clipboard

Challenge: Differential privacy is an important privacy concern when building statistical models on data containing sensitive information.
Approach: They propose a utility-preserving differentially private text transformation algorithm using auto-encoders that can be used to transform text to offer robustness against attacks and produce transformations with high semantic quality.
Outcome: The proposed model performs better against membership inference attacks while offering lower to no degradation in the utility of the underlying transformation process compared to baselines.
Controlling the Extraction of Memorized Data from Large Language Models via Prompt-Tuning (2023.acl-short)

Copied to clipboard

Challenge: Large Language Models memorize significant portions of training data, which poses privacy risk.
Approach: They propose a prompt-tuning approach to control the extraction rates of memorized content in large language models.
Outcome: The proposed techniques yield 9.3% increase in extraction rate compared to baseline model . the proposed defense achieves 97.7% reduction with a perplexity increase of 16.9% .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations