Papers with GDPR
Data Management Plan (DMP) for Language Data under the New General Da-ta Protection Regulation (GDPR) (L18-1)
Copied to clipboard
| Challenge: | ELRA proposes its own template for the Data Management Plan, which is being updated to take the new law into account. |
| Approach: | They propose a framework for the data management plan to be updated to take the new law into account and propose how it can be integrated into the DMP to increase transparency and spread good practices . |
| Outcome: | The proposed framework will strengthen certain principles related to the processing of personal data, which will also affect many projects in the field of natural language processing. |
Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning (2025.emnlp-main)
Copied to clipboard
Wenbin Hu, Haoran Li, Huihao Jing, Qi Hu, Ziqian Zeng, Sirui Han, Xu Heli, Tianshu Chu, Peizhao Hu, Yangqiu Song
| Challenge: | Current mitigation strategies fail to preserve contextual reasoning capabilities in risky scenarios, leading to systemic risks for legal compliance. |
| Approach: | They propose to use reinforcement learning with a rule-based reward to incentivize contextual reasoning capabilities while enhancing compliance with safety and privacy norms. |
| Outcome: | The proposed model outperforms Qwen2.5-7B-Instruct model in safety and privacy benchmarks and achieves +8.58% accuracy improvement. |
Lightweight Domain-Specific Language Model for Real-Time Structuring of Medical Prescriptions (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing language models ignore layout information, rely on expensive image-based architectures, or cannot operate under privacy and hardware constraints. |
| Approach: | They propose a lightweight, privacy-preserving transformer specifically designed for Entity Extraction (EE) and Entity Linking (EL) in french medical prescriptions. |
| Outcome: | The proposed model matches or surpasses larger document-understanding models on strict extraction metrics while maintaining essential spatial cues. |
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)
Copied to clipboard
| Challenge: | exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats. |
| Approach: | They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates. |
| Outcome: | The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data. |
C3PA: An Open Dataset of Expert-Annotated and Regulation-Aware Privacy Policies to Enable Scalable Regulatory Compliance Audits (2024.emnlp-main)
Copied to clipboard
Maaz Musa, Steven Winston, Garrison Allen, Jacob Schiller, Kevin Moore, Sean Quick, Johnathan Melvin, Padmini Srinivasan, Mihailis Diamantis, Rishab Nithyanand
| Challenge: | Privacy policies fall short of achieving compliance goals due to their inaccessibility or incomprehensibility. |
| Approach: | They propose to use C3PA to create an open regulation-aware dataset of expert-annotated privacy policies to aid automated audits of compliance with CCPA-related disclosure mandates. |
| Outcome: | The proposed dataset is uniquely suited for aiding automated audits of compliance with CCPA-related disclosure mandates from 411 unique organizations. |
ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance (2026.acl-long)
Copied to clipboard
Haoran Li, Yulin Chen, Huihao Jing, Wenbin Hu, Tsz Ho Li, Chanhou Lou, Hong Ting Tsang, Sirui Han, Yangqiu Song
| Challenge: | Existing approaches to contextualize safety and privacy assessments assume the availability of complete and clear context, whereas real-world contexts tend to be ambiguous and incomplete. |
| Approach: | They propose a semi-rule-based framework that leverages large language models to ground the input context in the legal domain and explicitly identify both known and unknown factors for legal compliance. |
| Outcome: | The proposed framework can significantly improve existing baselines without training and can identify the ambiguous and missing factors. |
Corpora of Disordered Speech in the Light of the GDPR: Two Use Cases from the DELAD Initiative (2020.lrec-1)
Copied to clipboard
| Challenge: | Corpora of disordered speech (CDS) are costly to collect and difficult to share due to personal data protection and IP issues. |
| Approach: | a new paper examines the legal grounds for processing corpora of disordered speech . it illustrates how consent and public interest are taken into consideration . the paper also examines how public interest research can be used to obtain consent . |
| Outcome: | a new study examines the legal grounds for processing corpora of disordered speech (CDS) two use cases illustrate the legal basis for processing CDS in light of the GDPR . |
Towards Operationalizing Right to Data Protection (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent work introduces the concept of generating unlearnable datasets (by adding imperceptible spurious correlations to the clean data) this approach is limited by several practical constraints like requiring knowledge of the target model. |
| Approach: | They propose a framework that injects imperceptible spurious correlations into natural language datasets, rendering them unlearnable without affecting semantic content. |
| Outcome: | The proposed framework can restrict newer models like GPT-4o and Llama from learning on generated data, resulting in a drop in test accuracy compared to their zero-shot performance. |
Privacy by Design and Language Resources (2020.lrec-1)
Copied to clipboard
| Challenge: | Privacy by Design is an approach in which privacy and data protection are embedded throughout the project lifecycle . the principle of Privacy by design was first mentioned in the 1995 EU Data Protection Directive . |
| Approach: | a paper proposes to analyze the practical meaning of Privacy by Design in the context of Language Resources . the paper propose measures and safeguards that can be implemented by the community to ensure respect of this principle. |
| Outcome: | The proposed paper analyzes the practical meaning of Privacy by Design in the context of Language Resources . proposed safeguards can be implemented by the community to ensure respect of this principle. |
S-RAG: A Novel Audit Framework for Detecting Unauthorized Use of Personal Data in RAG Systems (2025.acl-long)
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) systems rely on external data for accurate and context-specific responses. |
| Approach: | They propose a framework that enables users to determine whether their textual data has been utilized in RAG systems even in black-box settings with no prior system knowledge. |
| Outcome: | The proposed framework achieves an improvement in Accuracy by 19.9% while maintaining strong performance under adversarial defenses. |
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)
Copied to clipboard
Siddhant Arora, Henry Hosseini, Christine Utz, Vinayshekhar Bannihatti Kumar, Tristan Dhellemmes, Abhilasha Ravichander, Peter Story, Jasmine Mangat, Rex Chen, Martin Degeling, Thomas Norton, Thomas Hupperich, Shomir Wilson, Norman Sadeh
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
The DAPRECO Knowledge Base: Representing the GDPR in LegalRuleML (2020.lrec-1)
Copied to clipboard
| Challenge: | The DAPRECO knowledge base is a repository of rules written in LegalRuleML . the rules represent the provisions of the General Data Protection Regulation (GDPR) |
| Approach: | The DAPRECO knowledge base is a repository of rules written in LegalRuleML . the rules represent the provisions of the General Data Protection Regulation . |
| Outcome: | The DAPRECO knowledge base is the biggest knowledge base in LegalRuleML freely available online at (Robaldo et al., 2019). |
Adversarial Speech Generation and Natural Speech Recovery for Speech Content Protection (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, researchers focus on how to protect the speaker's identifiable information, represented as voiceprint, contained in the speech. |
| Approach: | They propose a frame-by-frame adversarial speech generation system to protect speech . they build an adversarials-based method that converts adversarially generated speech to human speech. |
| Outcome: | The proposed method can encode and recover any sensitive audio, and it is easy to be conducted with publicly available speech recognition technology. |
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation (2025.findings-acl)
Copied to clipboard
| Challenge: | Extensive experiments on public benchmarks and an in-house e-commerce dataset demonstrate Unilogit’s superior performance in balancing forget and retain objectives, outperforming state-of-the-art methods such as NPO and UnDIAL. |
| Approach: | They propose a self-distillation method that dynamically adjusts target logits to achieve a uniform probability for the target token. |
| Outcome: | Extensive experiments on public benchmarks and an in-house e-commerce dataset demonstrate Unilogit’s superior performance in balancing forget and retain objectives. |
LDEDE: LRP-Driven Efficient Detection and Editing Framework for LLM Privacy Neurons (2026.findings-acl)
Copied to clipboard
Zhao Zhengyuan, Cao Lifeng, null Sunhaodong, Shi Haotian, Du Xuehui, Liu Aodi, Niu Lanjie, Yang Xiaocheng
| Challenge: | Existing privacy protection methods fail to cover context-dependent sensitive information and are prone to performance degradation. |
| Approach: | They propose a Layer-wise Relevance Propagation-driven framework for efficient privacy neuron detection and editing. |
| Outcome: | The proposed framework achieves 80% higher efficiency than gradient attribution methods while reducing leakage risks of Phone, Email, and medical privacy by 42.7%–73.5% on average and cutting computational time by 60%–90%. |