Papers by Thomas Kleinbauer
Preventing Author Profiling through Zero-Shot Multilingual Back-Translation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Documents as short as a single sentence may reveal sensitive information about authors . style transfer is effective but a number of current methods cause a drop in down-stream utility . |
| Approach: | They propose a method to remove sensitive information from documents by multilingual back-translation using off-the-shelf translation models. |
| Outcome: | The proposed method lowers adversarial gender and race prediction by 22% while retaining 95% of original utility on downstream tasks. |
Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate Online (2022.lrec-1)
Copied to clipboard
Dana Ruiter, Liane Reiners, Ashwin Geet D’Sa, Thomas Kleinbauer, Dominique Fohr, Irina Illina, Dietrich Klakow, Christian Schemer, Angeliki Monnier
| Challenge: | HS-related corpora over-simplify the phenomenon of hate by labelling user content with binary classes, e.g., hate/neutral . this ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corporales. |
| Approach: | They present a corpus of 9k German and french user comments from migration-related news articles. |
| Outcome: | The proposed corpus is annotated with 23 features that become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate. |
Detection of Abusive Language: the Problem of Biased Datasets (N19-1)
Copied to clipboard
| Challenge: | Recent studies have reported high classification performance on datasets with difficult cases of abusive language. |
| Approach: | They examine the impact of data bias on abusive language detection by focusing on specific microposts rather than random sampling. |
| Outcome: | The proposed method is more accurate and more accurate than random sampling. |