Challenge: Privacy policies fall short of achieving compliance goals due to their inaccessibility or incomprehensibility.
Approach: They propose to use C3PA to create an open regulation-aware dataset of expert-annotated privacy policies to aid automated audits of compliance with CCPA-related disclosure mandates.
Outcome: The proposed dataset is uniquely suited for aiding automated audits of compliance with CCPA-related disclosure mandates from 411 unique organizations.

Similar Papers

A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant Identification (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that ignore law requirements are limited to English.
Approach: They construct a Chinese privacy policy dataset that can be used to analyze software privacy policies.
Outcome: The proposed dataset includes 483 Chinese Android privacy policies, over 11K sentences, and 52K fine-grained annotations.
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)

Copied to clipboard

Challenge: With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages.
Approach: They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA.
Outcome: The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices.
Building a Long Text Privacy Policy Corpus with Multi-Class Labels (2025.acl-long)

Copied to clipboard

Challenge: Legal text is susceptible to multiple valid, conflicting interpretations, and indeterminacy, interdependence between clauses, meaningful silence, and implications of legal defaults.
Approach: They propose to annotate privacy policies from 149 firms using a hand-coded dataset that captures key challenges peculiar to legal language.
Outcome: The proposed dataset includes privacy policies from 149 firms and includes materials incorporated by reference.
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation (2026.acl-long)

Copied to clipboard

Challenge: a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures.
Approach: They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English .
Outcome: The proposed framework outperforms large language models in terms of readability and accuracy.
Automated Detection and Analysis of Data Practices Using A Real-World Corpus (2024.findings-acl)

Copied to clipboard

Challenge: a crowd-sourced annotation tool matches data practices with policy excerpts . the complexity of privacy policies often deter users from reading them .
Approach: They propose an automated approach to identify and visualize data practices within privacy policies at different levels of detail.
Outcome: The proposed approach matches data practices with policy excerpts at different levels of detail.
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection .
Approach: They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions.
Outcome: The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions.
Question Answering for Privacy Policies: Combining Computational and Legal Perspectives (D19-1)

Copied to clipboard

Challenge: Privacy policies are long and complex documents that are difficult for users to read and understand.
Approach: They present a corpus of 1750 questions about privacy policies of mobile applications and over 3500 expert annotations of relevant answers.
Outcome: The proposed corpus of 1750 questions on privacy policies shows that a strong neural baseline underperforms human performance by almost 0.3 F1 on PrivacyQA.
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies (2021.acl-long)

Copied to clipboard

Challenge: Existing tools to interpret privacy policies have been used to understand them but there is a lack of large privacy policy corpora to simplify the process.
Approach: They propose to use a corpus of 1,005,380 English language privacy policies collected from the web to create semi-supervised and unsupervised models to interpret and simplify privacy policies.
Outcome: The proposed model outperforms all other publicly available privacy policy corpora and is ten times larger than the next largest public collection of privacy policies combined.
EROS:Entity-Driven Controlled Policy Document Summarization (2024.lrec-main)

Copied to clipboard

Challenge: a privacy policy is a crucial component of any organization that allows it to legally collect, process, store, and/or distribute personal data.
Approach: They propose to use a policy-document summarization dataset to enforce the summaries to include critical privacy-related entities and organization’s rationale in collecting those entities.
Outcome: The proposed model improves over baselines and qualitatively evaluates the proposed model on human and qualitative data.
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory (2025.naacl-long)

Copied to clipboard

Challenge: Existing privacy studies focus on sub-fields, but they focus on a few sub-domains.
Approach: They propose to use the Health Insurance Portability and Accountability Act of 1996 as an example to develop a checklist that covers social identities, private attributes, and existing privacy regulations.
Outcome: The proposed checklist covers social identities, private attributes, and existing privacy regulations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations