| Challenge: | Privacy by Design is an approach in which privacy and data protection are embedded throughout the project lifecycle . the principle of Privacy by design was first mentioned in the 1995 EU Data Protection Directive . |
| Approach: | a paper proposes to analyze the practical meaning of Privacy by Design in the context of Language Resources . the paper propose measures and safeguards that can be implemented by the community to ensure respect of this principle. |
| Outcome: | The proposed paper analyzes the practical meaning of Privacy by Design in the context of Language Resources . proposed safeguards can be implemented by the community to ensure respect of this principle. |
Similar Papers
Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy? (2021.acl-long)
Copied to clipboard
| Challenge: | Privacy policies are long and complex documents that are difficult for users to read and comprehend. |
| Approach: | They propose language technologies to help users reclaim control over their privacy . they highlight many remaining opportunities to develop more precise or nuanced language technologies . |
| Outcome: | The proposed language technologies can address the privacy information gap . they can be more precise or nuanced in the way they use the text of privacy policies. |
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)
Copied to clipboard
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, Norman Sadeh
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)
Copied to clipboard
Sonu Gupta, Geetika Gopi, Harish Balaji, Ellen Poplavska, Nora O’Toole, Siddhant Arora, Thomas Norton, Norman Sadeh, Shomir Wilson
| Challenge: | a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection . |
| Approach: | They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
| Outcome: | The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions. |
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are accessed via commercial APIs, but expose data to service providers. |
| Approach: | They propose a framework where a local model uses natural language instructions to rewrite queries and paired them with synthetic privacy profiles to achieve better privacy preservation. |
| Outcome: | The proposed model outperforms large-scale few-shot models in terms of privacy preservation and performance. |
Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy Policies (2021.acl-long)
Copied to clipboard
| Challenge: | Existing tools to interpret privacy policies have been used to understand them but there is a lack of large privacy policy corpora to simplify the process. |
| Approach: | They propose to use a corpus of 1,005,380 English language privacy policies collected from the web to create semi-supervised and unsupervised models to interpret and simplify privacy policies. |
| Outcome: | The proposed model outperforms all other publicly available privacy policy corpora and is ten times larger than the next largest public collection of privacy policies combined. |
Data Management Plan (DMP) for Language Data under the New General Da-ta Protection Regulation (GDPR) (L18-1)
Copied to clipboard
| Challenge: | ELRA proposes its own template for the Data Management Plan, which is being updated to take the new law into account. |
| Approach: | They propose a framework for the data management plan to be updated to take the new law into account and propose how it can be integrated into the DMP to increase transparency and spread good practices . |
| Outcome: | The proposed framework will strengthen certain principles related to the processing of personal data, which will also affect many projects in the field of natural language processing. |
Anonymisation Models for Text Data: State of the art, Challenges and Future Directions (2021.acl-long)
Copied to clipboard
| Challenge: | a paper examines the problem of automated text anonymisation . text anonymization is a prerequisite for secure sharing of documents containing sensitive information about individuals. |
| Approach: | They propose to incorporate explicit measures of disclosure risk into the text anonymisation process to reduce the risk of errors. |
| Outcome: | The proposed approach is based on a case study in which the authors outline the benefits and limitations of the proposed methods. |
The Model’s Language Matters: A Comparative Privacy Analysis of LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large language models are increasingly deployed in multilingual settings that process sensitive data . prior privacy evaluations focused on English, but new research shows that language matters for privacy leakage . |
| Approach: | They quantify six corpus-level linguistic indicators and evaluate vulnerability under three attack families. |
| Outcome: | The results show that language matters for privacy leakage in large language models . Italian exhibits the strongest exposure, while English and French are more resilient . |
Building a Long Text Privacy Policy Corpus with Multi-Class Labels (2025.acl-long)
Copied to clipboard
| Challenge: | Legal text is susceptible to multiple valid, conflicting interpretations, and indeterminacy, interdependence between clauses, meaningful silence, and implications of legal defaults. |
| Approach: | They propose to annotate privacy policies from 149 firms using a hand-coded dataset that captures key challenges peculiar to legal language. |
| Outcome: | The proposed dataset includes privacy policies from 149 firms and includes materials incorporated by reference. |
TextHide: Tackling Data Privacy in Language Understanding Tasks (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Unsolved privacy challenges in distributed or federated learning are a challenge for many domains including Natural Language Processing. |
| Approach: | They propose a federated learning framework that adds an encryption step to prevent an eavesdropping attacker from recovering private text data. |
| Outcome: | The proposed model can effectively defend against attacks on shared gradients or representations and the averaged accuracy reduction is only 1.9%. |