Papers by Elisabeth Eder
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection . |
| Approach: | They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse . |
| Outcome: | The proposed model could be improved to detect implicit abuse in a dataset with a standardized model. |
The Relevance of Value Systems for Offensive Language Detection (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent research in perspectivism has departed from the assumption that offensiveness can be defined through a universal perspective. |
| Approach: | They propose to use a dataset consisting of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems to identify offensiveness patterns. |
| Outcome: | The proposed dataset consists of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems. |
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)
Copied to clipboard
| Challenge: | exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats. |
| Approach: | They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates. |
| Outcome: | The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data. |
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)
Copied to clipboard
| Challenge: | lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems. |
| Approach: | They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation. |
| Outcome: | The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers. |
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)
Copied to clipboard
| Challenge: | Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web. |
| Approach: | They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances. |
| Outcome: | The proposed classifier augments training data with automatically-generated GPT-3 completions. |
CodE Alltag 2.0 — A Pseudonymized German-Language Email Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | unauthorized use of social media content as a data resource is often neglected . data privacy concerns are often overlooked in NLP research . |
| Approach: | They propose an algorithm for the protection of personal data via pseudonymization by automatically recognizing privacy-sensitive stretches of text in UGC. |
| Outcome: | The proposed algorithm protects personal data via pseudonymization on two hitherto non-anonymized German-language email corpora. |
Identifying Implicitly Abusive Remarks about Identity Groups using a Linguistically Informed Approach (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing datasets displaying high degree of implicit abuse are biased . current methods focus on explicit abuse, but there is little work on implicit forms of abuse . |
| Approach: | They propose to model atomic negative sentences to address implicit abuse by addressing its different subtypes and then separate them into subtype. |
| Outcome: | The proposed approach generalizes across different identities and languages. |
Beyond Negative Stereotypes – Non-Negative Abusive Utterances about Identity Groups and Their Semantic Variants (2025.acl-long)
Copied to clipboard
| Challenge: | implicitly abusive language is a language that could offend, demean or marginalize another person . a large portion of what is considered abusive language can be classified as implicitly abused . |
| Approach: | They propose to profile implicitly abusive language and use it to analyze a dataset of such utterances. |
| Outcome: | The proposed dataset identifies the type of abusive language that is not conveyed by unambiguously abusive words. |
A Question of Style: A Dataset for Analyzing Formality on Different Levels (2023.findings-eacl)
Copied to clipboard
| Challenge: | Using machine learning, we can produce contextually appropriate language. |
| Approach: | They present a dataset of German sentence-level formality assessed on a continuous informal-formal scale. |
| Outcome: | The proposed dataset compares sentences from a wide range of genres assessed on a continuous informal-formal scale. |