Papers by Ulrike Krieg-Holz

4 papers
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)

Copied to clipboard

Challenge: exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats.
Approach: They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates.
Outcome: The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data.
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)

Copied to clipboard

Challenge: lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems.
Approach: They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation.
Outcome: The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers.
CodE Alltag 2.0 — A Pseudonymized German-Language Email Corpus (2020.lrec-1)

Copied to clipboard

Challenge: unauthorized use of social media content as a data resource is often neglected . data privacy concerns are often overlooked in NLP research .
Approach: They propose an algorithm for the protection of personal data via pseudonymization by automatically recognizing privacy-sensitive stretches of text in UGC.
Outcome: The proposed algorithm protects personal data via pseudonymization on two hitherto non-anonymized German-language email corpora.
A Question of Style: A Dataset for Analyzing Formality on Different Levels (2023.findings-eacl)

Copied to clipboard

Challenge: Using machine learning, we can produce contextually appropriate language.
Approach: They present a dataset of German sentence-level formality assessed on a continuous informal-formal scale.
Outcome: The proposed dataset compares sentences from a wide range of genres assessed on a continuous informal-formal scale.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations