Papers by Jaydeep Borkar
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training (2025.findings-acl)
Copied to clipboard
Jaydeep Borkar, Matthew Jagielski, Katherine Lee, Niloofar Mireshghallah, David A. Smith, Christopher A. Choquette-Choo
| Challenge: | PII is a sensitive information that can be removed from large-language model training due to evolving curation techniques, or because it was recently scraped for retraining. |
| Approach: | They characterize a phenomenon where PII that appeared earlier in training becomes extractable at a later step after fine-tuning on other PI I. |
| Outcome: | The authors show that PII memorization is a dynamic property of a model that evolves throughout training pipelines and depends on commonly altered design choices. |