Papers by Kaushal Prajapati
De-Identification of Sensitive Personal Data in Datasets Derived from IIT-CDIP (2024.emnlp-main)
Copied to clipboard
Stefan Larson, Nicole Lima, Santiago Diaz, Amogh Joshi, Siddharth Betala, Jamiu Suleiman, Yash Mathur, Kaushal Prajapati, Ramla Alakraa, Junjie Shen, Temi Okotore, Kevin Leach
| Challenge: | Large volumes of data are becoming increasingly important for training machine learning models for document understanding tasks like classification, information extraction, and visual question answering. |
| Approach: | They propose a data de-identification pipeline that replaces sensitive data with synthetic, but realistic, data that preserves the utility of de-identified documents. |
| Outcome: | The proposed method preserves the utility of the de-identified documents so that they can continue to be used in various document understanding applications. |