| Challenge: | a recent study has raised concerns over privacy policies' opaqueness . lack of clarity in privacy policies can lead to undesired ads and privacy breaches . |
| Approach: | They propose to analyze the semantics of vague words and sentences and use them to identify vague content in privacy policies. |
| Outcome: | The proposed methods are effective and provide suggestions for improving privacy policies. |
Similar Papers
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools to interpret privacy policies have been used to understand them but there is a lack of large privacy policy corpora to simplify the process. |
| Approach: | They propose to use a corpus of 1,005,380 English language privacy policies collected from the web to create semi-supervised and unsupervised models to interpret and simplify privacy policies. |
| Outcome: | The proposed model outperforms all other publicly available privacy policy corpora and is ten times larger than the next largest public collection of privacy policies combined. |
Question Answering for Privacy Policies: Combining Computational and Legal Perspectives (D19-1)
Copied to clipboard
| Challenge: | Privacy policies are long and complex documents that are difficult for users to read and understand. |
| Approach: | They present a corpus of 1750 questions about privacy policies of mobile applications and over 3500 expert annotations of relevant answers. |
| Outcome: | The proposed corpus of 1750 questions on privacy policies shows that a strong neural baseline underperforms human performance by almost 0.3 F1 on PrivacyQA. |
Building a Long Text Privacy Policy Corpus with Multi-Class Labels (2025.acl-long)
Copied to clipboard
| Challenge: | Legal text is susceptible to multiple valid, conflicting interpretations, and indeterminacy, interdependence between clauses, meaningful silence, and implications of legal defaults. |
| Approach: | They propose to annotate privacy policies from 149 firms using a hand-coded dataset that captures key challenges peculiar to legal language. |
| Outcome: | The proposed dataset includes privacy policies from 149 firms and includes materials incorporated by reference. |
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation (2026.acl-long)
Copied to clipboard
Pengyun Zhu, Qiheng Sun, Long Wen, Yanbo Wang, Yang Cao, Junxu Liu, Deyi Xiong, Jinfei Liu, Zhibo Wang, Kui Ren
| Challenge: | a lack of high-quality English privacy policy corpus optimized for legal clarity and readability is limiting translation of privacy policies . 139 privacy policies are often considered "incomprehensible" due to technical jargon, legal language, and convoluted grammatical structures. |
| Approach: | They propose a high-quality English privacy policy corpus annotated by domain experts . they propose APPSI-139 to summarize and interpret privacy policies in English . |
| Outcome: | The proposed framework outperforms large language models in terms of readability and accuracy. |
Automated Detection and Analysis of Data Practices Using A Real-World Corpus (2024.findings-acl)
Copied to clipboard
| Challenge: | a crowd-sourced annotation tool matches data practices with policy excerpts . the complexity of privacy policies often deter users from reading them . |
| Approach: | They propose an automated approach to identify and visualize data practices within privacy policies at different levels of detail. |
| Outcome: | The proposed approach matches data practices with policy excerpts at different levels of detail. |
PLUE: Language Understanding Evaluation Benchmark for Privacy Policies in English (2023.acl-short)
Copied to clipboard
| Challenge: | Existing efforts to understand privacy policies are limited by processing the language in a way exclusive to a single task focusing on certain privacy practices. |
| Approach: | They propose a privacy policy language understanding evaluation benchmark to evaluate the understanding of privacy policies across multiple tasks. |
| Outcome: | The proposed framework improves the understanding of privacy policies across multiple tasks. |
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)
Copied to clipboard
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, Norman Sadeh
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
Intent Classification and Slot Filling for Privacy Policies (2021.acl-long)
Copied to clipboard
| Challenge: | Sentences written in privacy policies explain privacy practices and the constituent text spans convey further specific information. |
| Approach: | They propose an English corpus of 5,250 intent and 11,788 slot annotations . they propose two alternative neural approaches to model the corpus as a sequence-to-sequence learning task. |
| Outcome: | The proposed corpus predicts intent classification and slot filling, while the sequence tagging method outperforms slot filler by a large margin. |
Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy? (2021.acl-long)
Copied to clipboard
| Challenge: | Privacy policies are long and complex documents that are difficult for users to read and comprehend. |
| Approach: | They propose language technologies to help users reclaim control over their privacy . they highlight many remaining opportunities to develop more precise or nuanced language technologies . |
| Outcome: | The proposed language technologies can address the privacy information gap . they can be more precise or nuanced in the way they use the text of privacy policies. |
PolicyQA: A Reading Comprehension Dataset for Privacy Policies (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Privacy policy documents are long and verbose. Hence, a question answering system can help users find the information that is relevant and important to them. |
| Approach: | They propose to provide users with a short text span from policy documents to search for answers from a long text segment. |
| Outcome: | The proposed question answering system can help users find information relevant to them. |