A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)
Copied to clipboard
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, Norman Sadeh
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
Similar Papers
Building a Long Text Privacy Policy Corpus with Multi-Class Labels (2025.acl-long)
Copied to clipboard
| Challenge: | Legal text is susceptible to multiple valid, conflicting interpretations, and indeterminacy, interdependence between clauses, meaningful silence, and implications of legal defaults. |
| Approach: | They propose to annotate privacy policies from 149 firms using a hand-coded dataset that captures key challenges peculiar to legal language. |
| Outcome: | The proposed dataset includes privacy policies from 149 firms and includes materials incorporated by reference. |
A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant Identification (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets that ignore law requirements are limited to English. |
| Approach: | They construct a Chinese privacy policy dataset that can be used to analyze software privacy policies. |
| Outcome: | The proposed dataset includes 483 Chinese Android privacy policies, over 11K sentences, and 52K fine-grained annotations. |
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation (2026.acl-long)
Copied to clipboard
Pengyun Zhu, Qiheng Sun, Long Wen, Yanbo Wang, Yang Cao, Junxu Liu, Deyi Xiong, Jinfei Liu, Zhibo Wang, Kui Ren
| Challenge: | a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures. |
| Approach: | They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English . |
| Outcome: | The proposed framework outperforms large language models in terms of readability and accuracy. |
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)
Copied to clipboard
Sonu Gupta, Geetika Gopi, Harish Balaji, Ellen Poplavska, Nora O’Toole, Siddhant Arora, Thomas Norton, Norman Sadeh, Shomir Wilson
| Challenge: | a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection . |
| Approach: | They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
| Outcome: | The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy? (2021.acl-long)
Copied to clipboard
| Challenge: | Privacy policies are long and complex documents that are difficult for users to read and comprehend. |
| Approach: | They propose language technologies to help users reclaim control over their privacy . they highlight many remaining opportunities to develop more precise or nuanced language technologies . |
| Outcome: | The proposed language technologies can address the privacy information gap . they can be more precise or nuanced in the way they use the text of privacy policies. |
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools to interpret privacy policies have been used to understand them but there is a lack of large privacy policy corpora to simplify the process. |
| Approach: | They propose to use a corpus of 1,005,380 English language privacy policies collected from the web to create semi-supervised and unsupervised models to interpret and simplify privacy policies. |
| Outcome: | The proposed model outperforms all other publicly available privacy policy corpora and is ten times larger than the next largest public collection of privacy policies combined. |
C3PA: An Open Dataset of Expert-Annotated and Regulation-Aware Privacy Policies to Enable Scalable Regulatory Compliance Audits (2024.emnlp-main)
Copied to clipboard
Maaz Musa, Steven Winston, Garrison Allen, Jacob Schiller, Kevin Moore, Sean Quick, Johnathan Melvin, Padmini Srinivasan, Mihailis Diamantis, Rishab Nithyanand
| Challenge: | Privacy policies fall short of achieving compliance goals due to their inaccessibility or incomprehensibility. |
| Approach: | They propose to use C3PA to create an open regulation-aware dataset of expert-annotated privacy policies to aid automated audits of compliance with CCPA-related disclosure mandates. |
| Outcome: | The proposed dataset is uniquely suited for aiding automated audits of compliance with CCPA-related disclosure mandates from 411 unique organizations. |
PLUE: Language Understanding Evaluation Benchmark for Privacy Policies in English (2023.acl-short)
Copied to clipboard
| Challenge: | Existing efforts to understand privacy policies are limited by processing the language in a way exclusive to a single task focusing on certain privacy practices. |
| Approach: | They propose a privacy policy language understanding evaluation benchmark to evaluate the understanding of privacy policies across multiple tasks. |
| Outcome: | The proposed framework improves the understanding of privacy policies across multiple tasks. |
Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty . |
| Approach: | They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes . |
| Outcome: | The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies. |
Proceedings of the Second Workshop on Economics and Natural Language Processing (D19-51)
Copied to clipboard
| Challenge: | ECONLP 2019 will focus on the many ways natural language processing influences business relations and procedures . |
| Approach: | a talk will discuss use-cases of natural language processing to aid in regulatory workflows . a workshop will focus on the many ways how NLP influences business relations and procedures . |
| Outcome: | This talk covers use-cases of natural language processing to aid in regulatory workflows . it also discusses shortcomings of current NLP technologies for financial regulation . |