C3PA: An Open Dataset of Expert-Annotated and Regulation-Aware Privacy Policies to Enable Scalable Regulatory Compliance Audits (2024.emnlp-main)
Copied to clipboard
Maaz Musa, Steven Winston, Garrison Allen, Jacob Schiller, Kevin Moore, Sean Quick, Johnathan Melvin, Padmini Srinivasan, Mihailis Diamantis, Rishab Nithyanand
| Challenge: | Privacy policies fall short of achieving compliance goals due to their inaccessibility or incomprehensibility. |
| Approach: | They propose to use C3PA to create an open regulation-aware dataset of expert-annotated privacy policies to aid automated audits of compliance with CCPA-related disclosure mandates. |
| Outcome: | The proposed dataset is uniquely suited for aiding automated audits of compliance with CCPA-related disclosure mandates from 411 unique organizations. |
Similar Papers
A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant Identification (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that ignore law requirements are limited to English. |
| Approach: | They construct a Chinese privacy policy dataset that can be used to analyze software privacy policies. |
| Outcome: | The proposed dataset includes 483 Chinese Android privacy policies, over 11K sentences, and 52K fine-grained annotations. |
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)
Copied to clipboard
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, Norman Sadeh
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
Building a Long Text Privacy Policy Corpus with Multi-Class Labels (2025.acl-long)
Copied to clipboard
| Challenge: | Legal text is susceptible to multiple valid, conflicting interpretations, and indeterminacy, interdependence between clauses, meaningful silence, and implications of legal defaults. |
| Approach: | They propose to annotate privacy policies from 149 firms using a hand-coded dataset that captures key challenges peculiar to legal language. |
| Outcome: | The proposed dataset includes privacy policies from 149 firms and includes materials incorporated by reference. |
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation (2026.acl-long)
Copied to clipboard
Pengyun Zhu, Qiheng Sun, Long Wen, Yanbo Wang, Yang Cao, Junxu Liu, Deyi Xiong, Jinfei Liu, Zhibo Wang, Kui Ren
| Challenge: | a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures. |
| Approach: | They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English . |
| Outcome: | The proposed framework outperforms large language models in terms of readability and accuracy. |
Automated Detection and Analysis of Data Practices Using A Real-World Corpus (2024.findings-acl)
Copied to clipboard
| Challenge: | a crowd-sourced annotation tool matches data practices with policy excerpts . the complexity of privacy policies often deter users from reading them . |
| Approach: | They propose an automated approach to identify and visualize data practices within privacy policies at different levels of detail. |
| Outcome: | The proposed approach matches data practices with policy excerpts at different levels of detail. |
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)
Copied to clipboard
Sonu Gupta, Geetika Gopi, Harish Balaji, Ellen Poplavska, Nora O’Toole, Siddhant Arora, Thomas Norton, Norman Sadeh, Shomir Wilson
| Challenge: | a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection . |
| Approach: | They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
| Outcome: | The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
Question Answering for Privacy Policies: Combining Computational and Legal Perspectives (D19-1)
Copied to clipboard
| Challenge: | Privacy policies are long and complex documents that are difficult for users to read and understand. |
| Approach: | They present a corpus of 1750 questions about privacy policies of mobile applications and over 3500 expert annotations of relevant answers. |
| Outcome: | The proposed corpus of 1750 questions on privacy policies shows that a strong neural baseline underperforms human performance by almost 0.3 F1 on PrivacyQA. |
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools to interpret privacy policies have been used to understand them but there is a lack of large privacy policy corpora to simplify the process. |
| Approach: | They propose to use a corpus of 1,005,380 English language privacy policies collected from the web to create semi-supervised and unsupervised models to interpret and simplify privacy policies. |
| Outcome: | The proposed model outperforms all other publicly available privacy policy corpora and is ten times larger than the next largest public collection of privacy policies combined. |
EROS:Entity-Driven Controlled Policy Document Summarization (2024.lrec-main)
Copied to clipboard
| Challenge: | a privacy policy is a crucial component of any organization that allows it to legally collect, process, store, and/or distribute personal data. |
| Approach: | They propose to use a policy-document summarization dataset to enforce the summaries to include critical privacy-related entities and organization’s rationale in collecting those entities. |
| Outcome: | The proposed model improves over baselines and qualitatively evaluates the proposed model on human and qualitative data. |
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing privacy studies focus on sub-fields, but they focus on a few sub-domains. |
| Approach: | They propose to use the Health Insurance Portability and Accountability Act of 1996 as an example to develop a checklist that covers social identities, private attributes, and existing privacy regulations. |
| Outcome: | The proposed checklist covers social identities, private attributes, and existing privacy regulations. |