Papers by Kaushal Prajapati

1 papers
De-Identification of Sensitive Personal Data in Datasets Derived from IIT-CDIP (2024.emnlp-main)

Copied to clipboard

Challenge: Large volumes of data are becoming increasingly important for training machine learning models for document understanding tasks like classification, information extraction, and visual question answering.
Approach: They propose a data de-identification pipeline that replaces sensitive data with synthetic, but realistic, data that preserves the utility of de-identified documents.
Outcome: The proposed method preserves the utility of the de-identified documents so that they can continue to be used in various document understanding applications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations