Challenge: With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages.
Approach: They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA.
Outcome: The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices.

Similar Papers

Building a Long Text Privacy Policy Corpus with Multi-Class Labels (2025.acl-long)

Copied to clipboard

Challenge: Legal text is susceptible to multiple valid, conflicting interpretations, and indeterminacy, interdependence between clauses, meaningful silence, and implications of legal defaults.
Approach: They propose to annotate privacy policies from 149 firms using a hand-coded dataset that captures key challenges peculiar to legal language.
Outcome: The proposed dataset includes privacy policies from 149 firms and includes materials incorporated by reference.
A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant Identification (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets that ignore law requirements are limited to English.
Approach: They construct a Chinese privacy policy dataset that can be used to analyze software privacy policies.
Outcome: The proposed dataset includes 483 Chinese Android privacy policies, over 11K sentences, and 52K fine-grained annotations.
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation (2026.acl-long)

Copied to clipboard

Challenge: a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures.
Approach: They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English .
Outcome: The proposed framework outperforms large language models in terms of readability and accuracy.
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection .
Approach: They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions.
Outcome: The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions.
Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy? (2021.acl-long)

Copied to clipboard

Challenge: Privacy policies are long and complex documents that are difficult for users to read and comprehend.
Approach: They propose language technologies to help users reclaim control over their privacy . they highlight many remaining opportunities to develop more precise or nuanced language technologies .
Outcome: The proposed language technologies can address the privacy information gap . they can be more precise or nuanced in the way they use the text of privacy policies.
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies (2021.acl-long)

Copied to clipboard

Challenge: Existing tools to interpret privacy policies have been used to understand them but there is a lack of large privacy policy corpora to simplify the process.
Approach: They propose to use a corpus of 1,005,380 English language privacy policies collected from the web to create semi-supervised and unsupervised models to interpret and simplify privacy policies.
Outcome: The proposed model outperforms all other publicly available privacy policy corpora and is ten times larger than the next largest public collection of privacy policies combined.
C3PA: An Open Dataset of Expert-Annotated and Regulation-Aware Privacy Policies to Enable Scalable Regulatory Compliance Audits (2024.emnlp-main)

Copied to clipboard

Challenge: Privacy policies fall short of achieving compliance goals due to their inaccessibility or incomprehensibility.
Approach: They propose to use C3PA to create an open regulation-aware dataset of expert-annotated privacy policies to aid automated audits of compliance with CCPA-related disclosure mandates.
Outcome: The proposed dataset is uniquely suited for aiding automated audits of compliance with CCPA-related disclosure mandates from 411 unique organizations.
PLUE: Language Understanding Evaluation Benchmark for Privacy Policies in English (2023.acl-short)

Copied to clipboard

Challenge: Existing efforts to understand privacy policies are limited by processing the language in a way exclusive to a single task focusing on certain privacy practices.
Approach: They propose a privacy policy language understanding evaluation benchmark to evaluate the understanding of privacy policies across multiple tasks.
Outcome: The proposed framework improves the understanding of privacy policies across multiple tasks.
Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty .
Approach: They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes .
Outcome: The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies.
Proceedings of the Second Workshop on Economics and Natural Language Processing (D19-51)

Copied to clipboard

Challenge: ECONLP 2019 will focus on the many ways natural language processing influences business relations and procedures .
Approach: a talk will discuss use-cases of natural language processing to aid in regulatory workflows . a workshop will focus on the many ways how NLP influences business relations and procedures .
Outcome: This talk covers use-cases of natural language processing to aid in regulatory workflows . it also discusses shortcomings of current NLP technologies for financial regulation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations