Papers by Michael Wiegand

21 papers
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)

Copied to clipboard

Challenge: Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection .
Approach: They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse .
Outcome: The proposed model could be improved to detect implicit abuse in a dataset with a standardized model.
The Relevance of Value Systems for Offensive Language Detection (2026.eacl-long)

Copied to clipboard

Challenge: Recent research in perspectivism has departed from the assumption that offensiveness can be defined through a universal perspective.
Approach: They propose to use a dataset consisting of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems to identify offensiveness patterns.
Outcome: The proposed dataset consists of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems.
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)

Copied to clipboard

Challenge: exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats.
Approach: They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates.
Outcome: The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data.
Revisiting Implicitly Abusive Language Detection: Evaluating LLMs in Zero-Shot and Few-Shot Settings (2025.coling-main)

Copied to clipboard

Challenge: Current research focuses on explicit abusive language, but subtler forms of IAL remain insufficiently studied.
Approach: They evaluate the models' capabilities in classifying sentences directly as either IAL or benign, and in extracting linguistic features associated with IAL.
Outcome: The proposed models outperform the best previously reported methods in classifying sentences directly as IAL or benign and extracting linguistic features associated with IAL.
Implicitly Abusive Comparisons – A New Dataset and Linguistic Analysis (2021.eacl-main)

Copied to clipboard

Challenge: Using crowdsourcing, we can detect implicitly abusive comparisons . Abusive language is defined as hurtful, derogatory or obscene utterances made by one person to another .
Approach: They propose to use crowdsourcing to generate a dataset for detecting implicitly abusive comparisons . they also use a range of linguistic features to better understand abusive comparison mechanisms .
Outcome: The proposed dataset includes measures to obtain representative and unbiased comparisons.
Distinguishing affixoid formations from compounds (C18-1)

Copied to clipboard

Challenge: affixoids are morphemes in between affids and free stems that have been associated with increased productivity and a bleached semantics but not empirically validated.
Approach: They propose to use affixoids as morphemes in between affids and stems to test their classification in a subset of German words that includes many hapaxes.
Outcome: The proposed morpheme can be classed as affixoid or non-affixoids with a best F1 score of 74% on a subset of German words that includes many hapaxes .
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web.
Approach: They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances.
Outcome: The proposed classifier augments training data with automatically-generated GPT-3 completions.
Inducing a Lexicon of Abusive Words – a Feature-Based Approach (N18-1)

Copied to clipboard

Challenge: a new classification task is needed to identify abusive words among a set of negative polar expressions.
Approach: They propose to calibrate a domain-independent lexicon for detection of abusive words . they use a small manually annotated base lexico to calibrated a large lexical .
Outcome: The proposed feature can be calibrated on a small manually annotated base lexicon and produced on large datasets.
Automatically Creating a Lexicon of Verbal Polarity Shifters: Mono- and Cross-lingual Methods for German (C18-1)

Copied to clipboard

Challenge: a large number of verbal polarity shifters are available for multiple languages, but only English has a sizable lexicon of them.
Approach: They use methods to create large lexicon of verbal polarity shifters in germany . they bootstrap annotated verbs with a supervised classifier and apply them to German .
Outcome: The proposed method is able to create a large lexicon of verbal polarity shifters in germany . it reduces annotation effort by leveraging cross-lingual information from the English lexico .
Disambiguation of Verbal Shifters (L18-1)

Copied to clipboard

Challenge: Negation is a contextual phenomenon that needs to be addressed in sentiment analysis.
Approach: They propose a supervised learning approach to disambiguate verbal shifters using generalization features and a new lexicon.
Outcome: The proposed approach takes into account various features, particularly generalization features.
Introducing a Lexicon of Verbal Polarity Shifters for English (L18-1)

Copied to clipboard

Challenge: Negation words can change the sentiment polarity of a phrase, but there are more than 1200 other polarities.
Approach: They propose a lexicon of verbal polarity shifters that covers the entirety of verbs found in WordNet.
Outcome: The proposed lexicon covers the entirety of verbs found in WordNet.
Identifying Implicitly Abusive Remarks about Identity Groups using a Linguistically Informed Approach (2022.naacl-main)

Copied to clipboard

Challenge: Existing datasets displaying high degree of implicit abuse are biased . current methods focus on explicit abuse, but there is little work on implicit forms of abuse .
Approach: They propose to model atomic negative sentences to address implicit abuse by addressing its different subtypes and then separate them into subtype.
Outcome: The proposed approach generalizes across different identities and languages.
Beyond Negative Stereotypes – Non-Negative Abusive Utterances about Identity Groups and Their Semantic Variants (2025.acl-long)

Copied to clipboard

Challenge: implicitly abusive language is a language that could offend, demean or marginalize another person . a large portion of what is considered abusive language can be classified as implicitly abused .
Approach: They propose to profile implicitly abusive language and use it to analyze a dataset of such utterances.
Outcome: The proposed dataset identifies the type of abusive language that is not conveyed by unambiguously abusive words.
A Question of Style: A Dataset for Analyzing Formality on Different Levels (2023.findings-eacl)

Copied to clipboard

Challenge: Using machine learning, we can produce contextually appropriate language.
Approach: They present a dataset of German sentence-level formality assessed on a continuous informal-formal scale.
Outcome: The proposed dataset compares sentences from a wide range of genres assessed on a continuous informal-formal scale.
Oddballs and Misfits: Detecting Implicit Abuse in Which Identity Groups are Depicted as Deviating from the Norm (2024.emnlp-main)

Copied to clipboard

Challenge: Abusive language is often defined as hurtful, derogatory or obscene utterances made by one person to another.
Approach: They propose to use a dataset to detect abusive sentences in identity groups . they also report on classification experiments.
Outcome: The proposed dataset includes 7 identity groups and includes classification experiments.
Exploiting Emojis for Abusive Language Detection (2021.eacl-main)

Copied to clipboard

Challenge: emojis can be used as a proxy for learning a lexicon of abusive words . eliot safina and samuel khan are the authors of this paper .
Approach: They propose to use abusive emojis as a proxy for learning a lexicon of abusive words.
Outcome: The proposed approach generates a lexicon that performs as well as the most advanced lexical induction method.
Biographically Relevant Tweets – a New Dataset, Linguistic Analysis and Classification Experiments (2022.coling-1)

Copied to clipboard

Challenge: Unlike previous work, we do not restrict biographical relevance to a small fixed set of pre-defined relations.
Approach: They propose a dataset comprising tweets for the novel task of detecting biographically relevant utterances.
Outcome: The proposed dataset focuses on biographical information on ordinary users of Twitter.
Detecting Derogatory Compounds – An Unsupervised Approach (N19-1)

Copied to clipboard

Challenge: Derogatory compounds are more difficult to detect than derogatory unigrams since they are sparsely represented in general-purpose lexical resources.
Approach: They propose an unsupervised classification approach that incorporates linguistic properties of compounds.
Outcome: The proposed method is compared with existing methods for extracting derogatory unigrams . the proposed method uses a distributional representation to incorporate linguistic properties of compounds .
Doctor Who? Framing Through Names and Titles in German (2020.lrec-1)

Copied to clipboard

Challenge: Entity framing is the selection of aspects of an entity to promote a particular viewpoint towards that entity.
Approach: They investigate entity framing of political figures through the use of names and titles in German online discourse.
Outcome: The proposed method improves existing studies on German political discourse . it shows that the formality of naming correlates positively with stance in the tweets .
Enhancing a Lexicon of Polarity Shifters through the Supervised Classification of Shifting Directions (2020.lrec-1)

Copied to clipboard

Challenge: Existing polarity shifter lexica only specify when a word can cause shifting, but do not specify when this is limited to a single shifting direction.
Approach: They propose a classifier that determines the shifting direction of polarity shifters by using resource-driven features and data-driven feature.
Outcome: The proposed classifier enhances the largest available polarity shifter lexicon.
Detection of Abusive Language: the Problem of Biased Datasets (N19-1)

Copied to clipboard

Challenge: Recent studies have reported high classification performance on datasets with difficult cases of abusive language.
Approach: They examine the impact of data bias on abusive language detection by focusing on specific microposts rather than random sampling.
Outcome: The proposed method is more accurate and more accurate than random sampling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations