Challenge: unauthorized use of social media content as a data resource is often neglected . data privacy concerns are often overlooked in NLP research .
Approach: They propose an algorithm for the protection of personal data via pseudonymization by automatically recognizing privacy-sensitive stretches of text in UGC.
Outcome: The proposed algorithm protects personal data via pseudonymization on two hitherto non-anonymized German-language email corpora.

Similar Papers

“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)

Copied to clipboard

Challenge: exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats.
Approach: They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates.
Outcome: The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)

Copied to clipboard

Challenge: EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
Approach: They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics.
Outcome: The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data.
Sharing Copies of Synthetic Clinical Corpora without Physical Distribution — A Case Study to Get Around IPRs and Privacy Constraints Featuring the German JSYNCC Corpus (L18-1)

Copied to clipboard

Challenge: eu legal culture imposes unsurmountable hurdles to exploit copyright protected language data . legal constraints have seriously hampered progress in resource-greedy NLP research . authors propose a new approach for the creation and re-use of clinical corpora .
Approach: They propose a method for the creation and re-use of clinical corpora based on a two-step workflow . they substitute authentic clinical documents by synthetic ones, i.e., made-up reports and case studies .
Outcome: a new approach replaces authentic clinical documents by synthetic ones, i.e., made-up reports and case studies published in medical e-textbooks.
Building a Corpus from Handwritten Picture Postcards: Transcription, Annotation and Part-of-Speech Tagging (L18-1)

Copied to clipboard

Challenge: In this paper, we describe the processes and challenges of digitalisation, manual transcription, and manual annotation of over 11,000 postcards.
Approach: They describe the processes and challenges of digitalisation, manual transcription, and manual annotation of over 11,000 postcards written in German and Swiss German.
Outcome: The proposed system outperforms state-of-the-art taggers in the evaluation of the 'picture postcard corpus' containing over 11,000 handwritten postcards .
Privacy-Preserving Natural Language Processing (2023.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial will help the NLP community to get familiar with current research in privacy-preserving methods.
Approach: This tutorial will help the NLP community to get familiar with current research in privacy-preserving methods.
Outcome: The tutorial will cover membership inference, differential privacy, homomorphic encryption, or federated learning, all with typical use-cases and potential pitfalls.
CyberAgressionAdo-v2: Leveraging Pragmatic-Level Information to Decipher Online Hate in French Multiparty Chats (2024.lrec-main)

Copied to clipboard

Challenge: Using a hierarchical tagset, cyberbullying narratives are described in the dataset CyberAgressionAdo-V1 . resulting dataset comprises 19 conversations that have been manually annotated .
Approach: They propose a new tagset that includes tags marking pragmatic-level information occurring in cyberbullying situations.
Outcome: The proposed tagset includes tags marking pragmatic-level information occurring in cyberbullying situations.
Pseudonymization Categories across Domain Boundaries (2024.lrec-main)

Copied to clipboard

Challenge: Linguistic data can contain personal information, which is limited in accessibility . a universal system of tags for categorizing PIIs could be developed to replace them .
Approach: They analyze tagsets used for anonymization and pseudonymization to find out what kinds of PII appear in different domains.
Outcome: The proposed system would allow for dynamic pseudonymization while keeping the data readable and useful for future research.
SOBR: A Corpus for Stylometry, Obfuscation, and Bias on Reddit (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora are limited in scope and can be used to collect data on author attributes.
Approach: They propose to use subreddits, flairs, and self-reports as distant labels for author attributes (age, gender, nationality, personality, and political leaning) .
Outcome: The proposed method could be used to infer author attributes from public posts despite their discreetness and anonymity .
Towards Privacy by Design in Learner Corpora Research: A Case of On-the-fly Pseudonymization of Swedish Learner Essays (2020.coling-main)

Copied to clipboard

Challenge: An ongoing project aims at automating pseudonymization of learner essays . 89% of the personal information can be successfully identified in learner data .
Approach: They propose to use rule-based methods to detect 15 categories out of 19 suggested by the authors.
Outcome: The proposed methods detect 15 categories out of 19 suggested by the authors . 89% of the personal information can be successfully identified in learner data and annotated correctly with an inter-annotator agreement of 86% measured as Fleiss kappa and Krippendorff’s alpha.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations