| Challenge: | unauthorized use of social media content as a data resource is often neglected . data privacy concerns are often overlooked in NLP research . |
| Approach: | They propose an algorithm for the protection of personal data via pseudonymization by automatically recognizing privacy-sensitive stretches of text in UGC. |
| Outcome: | The proposed algorithm protects personal data via pseudonymization on two hitherto non-anonymized German-language email corpora. |
Similar Papers
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)
Copied to clipboard
| Challenge: | exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats. |
| Approach: | They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates. |
| Outcome: | The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
Sharing Copies of Synthetic Clinical Corpora without Physical Distribution — A Case Study to Get Around IPRs and Privacy Constraints Featuring the German JSYNCC Corpus (L18-1)
Copied to clipboard
| Challenge: | eu legal culture imposes unsurmountable hurdles to exploit copyright protected language data . legal constraints have seriously hampered progress in resource-greedy NLP research . authors propose a new approach for the creation and re-use of clinical corpora . |
| Approach: | They propose a method for the creation and re-use of clinical corpora based on a two-step workflow . they substitute authentic clinical documents by synthetic ones, i.e., made-up reports and case studies . |
| Outcome: | a new approach replaces authentic clinical documents by synthetic ones, i.e., made-up reports and case studies published in medical e-textbooks. |
Building a Corpus from Handwritten Picture Postcards: Transcription, Annotation and Part-of-Speech Tagging (L18-1)
Copied to clipboard
| Challenge: | In this paper, we describe the processes and challenges of digitalisation, manual transcription, and manual annotation of over 11,000 postcards. |
| Approach: | They describe the processes and challenges of digitalisation, manual transcription, and manual annotation of over 11,000 postcards written in German and Swiss German. |
| Outcome: | The proposed system outperforms state-of-the-art taggers in the evaluation of the 'picture postcard corpus' containing over 11,000 handwritten postcards . |
Privacy-Preserving Natural Language Processing (2023.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial will help the NLP community to get familiar with current research in privacy-preserving methods. |
| Approach: | This tutorial will help the NLP community to get familiar with current research in privacy-preserving methods. |
| Outcome: | The tutorial will cover membership inference, differential privacy, homomorphic encryption, or federated learning, all with typical use-cases and potential pitfalls. |
CyberAgressionAdo-v2: Leveraging Pragmatic-Level Information to Decipher Online Hate in French Multiparty Chats (2024.lrec-main)
Copied to clipboard
| Challenge: | Using a hierarchical tagset, cyberbullying narratives are described in the dataset CyberAgressionAdo-V1 . resulting dataset comprises 19 conversations that have been manually annotated . |
| Approach: | They propose a new tagset that includes tags marking pragmatic-level information occurring in cyberbullying situations. |
| Outcome: | The proposed tagset includes tags marking pragmatic-level information occurring in cyberbullying situations. |
Pseudonymization Categories across Domain Boundaries (2024.lrec-main)
Copied to clipboard
Maria Irena Szawerna, Simon Dobnik, Therese Lindström Tiedemann, Ricardo Muñoz Sánchez, Xuan-Son Vu, Elena Volodina
| Challenge: | Linguistic data can contain personal information, which is limited in accessibility . a universal system of tags for categorizing PIIs could be developed to replace them . |
| Approach: | They analyze tagsets used for anonymization and pseudonymization to find out what kinds of PII appear in different domains. |
| Outcome: | The proposed system would allow for dynamic pseudonymization while keeping the data readable and useful for future research. |
SOBR: A Corpus for Stylometry, Obfuscation, and Bias on Reddit (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing corpora are limited in scope and can be used to collect data on author attributes. |
| Approach: | They propose to use subreddits, flairs, and self-reports as distant labels for author attributes (age, gender, nationality, personality, and political leaning) . |
| Outcome: | The proposed method could be used to infer author attributes from public posts despite their discreetness and anonymity . |
Towards Privacy by Design in Learner Corpora Research: A Case of On-the-fly Pseudonymization of Swedish Learner Essays (2020.coling-main)
Copied to clipboard
| Challenge: | An ongoing project aims at automating pseudonymization of learner essays . 89% of the personal information can be successfully identified in learner data . |
| Approach: | They propose to use rule-based methods to detect 15 categories out of 19 suggested by the authors. |
| Outcome: | The proposed methods detect 15 categories out of 19 suggested by the authors . 89% of the personal information can be successfully identified in learner data and annotated correctly with an inter-annotator agreement of 86% measured as Fleiss kappa and Krippendorff’s alpha. |