Papers by Josef Ruppenhofer
Implicitly Abusive Language – What does it actually look like and why are we not getting there? (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing datasets make learning implicit abuse difficult, argues a new position paper . a lack of work on implicit abuse has limited the effectiveness of automatic detection . |
| Approach: | They argue that existing datasets make learning implicit abuse difficult . they propose a divide-and-conquer strategy to detect implicit abuse . |
| Outcome: | The proposed model could be improved to detect implicit abuse in a dataset with a standardized model. |
The Relevance of Value Systems for Offensive Language Detection (2026.eacl-long)
Copied to clipboard
| Challenge: | Recent research in perspectivism has departed from the assumption that offensiveness can be defined through a universal perspective. |
| Approach: | They propose to use a dataset consisting of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems to identify offensiveness patterns. |
| Outcome: | The proposed dataset consists of neutrally-phrased sentences on controversial topics, evaluated by individuals from 4 different value systems. |
Out of the Mouths of MPs: Speaker Attribution in Parliamentary Debates (2024.lrec-main)
Copied to clipboard
| Challenge: | Identifying who says what to whom is an essential prerequisite for analysing human communication. |
| Approach: | They propose a new corpus for speaker attribution in german parliamentary debates . the data includes more than 7,700 manually annotated events of speech, thought and writing . they then apply their model to predict speech events in 20 years of debates and investigate the use of factives in the rhetoric of MPs. |
| Outcome: | The proposed model predicts speech events in 20 years of debates and investigates the use of factives in the rhetoric of MPs. |
Implicitly Abusive Comparisons – A New Dataset and Linguistic Analysis (2021.eacl-main)
Copied to clipboard
| Challenge: | Using crowdsourcing, we can detect implicitly abusive comparisons . Abusive language is defined as hurtful, derogatory or obscene utterances made by one person to another . |
| Approach: | They propose to use crowdsourcing to generate a dataset for detecting implicitly abusive comparisons . they also use a range of linguistic features to better understand abusive comparison mechanisms . |
| Outcome: | The proposed dataset includes measures to obtain representative and unbiased comparisons. |
Distinguishing affixoid formations from compounds (C18-1)
Copied to clipboard
| Challenge: | affixoids are morphemes in between affids and free stems that have been associated with increased productivity and a bleached semantics but not empirically validated. |
| Approach: | They propose to use affixoids as morphemes in between affids and stems to test their classification in a subset of German words that includes many hapaxes. |
| Outcome: | The proposed morpheme can be classed as affixoid or non-affixoids with a best F1 score of 74% on a subset of German words that includes many hapaxes . |
Who’s in, who’s out? Predicting the Inclusiveness or Exclusiveness of Personal Pronouns in Parliamentary Debates (2022.lrec-1)
Copied to clipboard
| Challenge: | clusivity properties of personal pronouns are captured in context, including/excluding audience and/or non-speech act participants. |
| Approach: | They propose a compositional annotation scheme to capture the clusivity properties of personal pronouns in context, which is their ability to construct and manage in-groups and out-group. |
| Outcome: | The proposed schema achieves high inter-annotator agreement with a Cohen’s in the range of 89.7-93.2 and a percentage agreement of > 96%. |
Euphemistic Abuse – A New Dataset and Classification Experiments for Implicitly Abusive Language (2023.emnlp-main)
Copied to clipboard
| Challenge: | Currently, only explicit abuse can be reliably detected due to the increasing amount of abusive language on the Web. |
| Approach: | They propose a crowdsourced dataset that can detect euphemistic abuse by paraphrasing simple explicit utterances. |
| Outcome: | The proposed classifier augments training data with automatically-generated GPT-3 completions. |
Every Verb in Its Right Place? A Roadmap for Operationalizing Developmental Stages in the Acquisition of L2 German (2024.lrec-main)
Copied to clipboard
| Challenge: | Developmental stages are a linguistic concept claiming that language learning progresses in an ordered, step-like manner. |
| Approach: | They propose to translate a linguistic specification into a computational procedure that can assign clauses to a developmental stage based on verb placement. |
| Outcome: | The proposed system lacks a coherent linguistic specification of developmental stages . it also lacks the ability to translate the specification into a computational procedure based on verb placement. |
Inducing a Lexicon of Abusive Words – a Feature-Based Approach (N18-1)
Copied to clipboard
| Challenge: | a new classification task is needed to identify abusive words among a set of negative polar expressions. |
| Approach: | They propose to calibrate a domain-independent lexicon for detection of abusive words . they use a small manually annotated base lexico to calibrated a large lexical . |
| Outcome: | The proposed feature can be calibrated on a small manually annotated base lexicon and produced on large datasets. |
Automatically Creating a Lexicon of Verbal Polarity Shifters: Mono- and Cross-lingual Methods for German (C18-1)
Copied to clipboard
| Challenge: | a large number of verbal polarity shifters are available for multiple languages, but only English has a sizable lexicon of them. |
| Approach: | They use methods to create large lexicon of verbal polarity shifters in germany . they bootstrap annotated verbs with a supervised classifier and apply them to German . |
| Outcome: | The proposed method is able to create a large lexicon of verbal polarity shifters in germany . it reduces annotation effort by leveraging cross-lingual information from the English lexico . |
Disambiguation of Verbal Shifters (L18-1)
Copied to clipboard
| Challenge: | Negation is a contextual phenomenon that needs to be addressed in sentiment analysis. |
| Approach: | They propose a supervised learning approach to disambiguate verbal shifters using generalization features and a new lexicon. |
| Outcome: | The proposed approach takes into account various features, particularly generalization features. |
Introducing a Lexicon of Verbal Polarity Shifters for English (L18-1)
Copied to clipboard
| Challenge: | Negation words can change the sentiment polarity of a phrase, but there are more than 1200 other polarities. |
| Approach: | They propose a lexicon of verbal polarity shifters that covers the entirety of verbs found in WordNet. |
| Outcome: | The proposed lexicon covers the entirety of verbs found in WordNet. |
Identifying Implicitly Abusive Remarks about Identity Groups using a Linguistically Informed Approach (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing datasets displaying high degree of implicit abuse are biased . current methods focus on explicit abuse, but there is little work on implicit forms of abuse . |
| Approach: | They propose to model atomic negative sentences to address implicit abuse by addressing its different subtypes and then separate them into subtype. |
| Outcome: | The proposed approach generalizes across different identities and languages. |
Sprucing up the trees – Error detection in treebanks (C18-1)
Copied to clipboard
| Challenge: | a method for detecting annotation errors in manually annotated dependency trees is presented . the method is based on ensemble parsing and Bayesian inference guided by active learning . |
| Approach: | They propose a method for detecting annotation errors in manually annotated dependency parse trees . they use ensemble parsing in combination with Bayesian inference guided by active learning . |
| Outcome: | The proposed method detects errors in annotated dependency treebanks and improves parsing accuracy on in- and out-of-domain data. |
Beyond Negative Stereotypes – Non-Negative Abusive Utterances about Identity Groups and Their Semantic Variants (2025.acl-long)
Copied to clipboard
| Challenge: | implicitly abusive language is a language that could offend, demean or marginalize another person . a large portion of what is considered abusive language can be classified as implicitly abused . |
| Approach: | They propose to profile implicitly abusive language and use it to analyze a dataset of such utterances. |
| Outcome: | The proposed dataset identifies the type of abusive language that is not conveyed by unambiguously abusive words. |
Fine-grained Named Entity Annotations for German Biographic Interviews (2020.lrec-1)
Copied to clipboard
| Challenge: | a NER annotation scheme is adapted for a corpus of transcripts of biographic interviews with emigrants to German . a dataset of spoken data and teaser tweets from newspaper sites are used to test the NER inventory. |
| Approach: | They propose a fine-grained NER annotation scheme with 30 labels and apply it to German data. |
| Outcome: | The proposed NER annotations can be applied to spoken data and teaser tweets from newspaper sites and achieve good inter-annotator agreement. |
A New Resource for German Causal Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Annotations of causal language are challenging for automatic and human annotators. |
| Approach: | They propose a German causal annotation resource with annotations in context for verbs, nouns and prepositions. |
| Outcome: | The proposed annotation scheme distinguishes three types of causal events . the proposed framework also provides annotations for semantic roles and actors . |
Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies (2020.lrec-1)
Copied to clipboard
Manuela Sanguinetti, Cristina Bosco, Lauren Cassidy, Özlem Çetinoğlu, Alessandra Teresa Cignarella, Teresa Lynn, Ines Rehbein, Josef Ruppenhofer, Djamé Seddah, Amir Zeldes
| Challenge: | Despite the increasing number of contributions on Part-of-Speech tagging and parsing, automatic processing of user-generated content (UGC) still represents a challenging task. |
| Approach: | They propose a set of guidelines for the annotation of user-generated texts within the Universal Dependencies framework. |
| Outcome: | The proposed annotation guidelines promote cross-linguistic consistency, which has always been in the spirit of UD. |
Building a Morphological Treebank for German from a Linguistic Database (L18-1)
Copied to clipboard
| Challenge: | German is a language with complex morphological processes. |
| Approach: | They propose a morphological treebank for German based on a German morphology database and a Perl script for the generation. |
| Outcome: | The proposed treebank is based on the German lexical database CELEX and is able to generate 40,000 morphological trees with a grade of detail that can be chosen according to the requirements of the applications. |
Oddballs and Misfits: Detecting Implicit Abuse in Which Identity Groups are Depicted as Deviating from the Norm (2024.emnlp-main)
Copied to clipboard
| Challenge: | Abusive language is often defined as hurtful, derogatory or obscene utterances made by one person to another. |
| Approach: | They propose to use a dataset to detect abusive sentences in identity groups . they also report on classification experiments. |
| Outcome: | The proposed dataset includes 7 identity groups and includes classification experiments. |
Exploiting Emojis for Abusive Language Detection (2021.eacl-main)
Copied to clipboard
| Challenge: | emojis can be used as a proxy for learning a lexicon of abusive words . eliot safina and samuel khan are the authors of this paper . |
| Approach: | They propose to use abusive emojis as a proxy for learning a lexicon of abusive words. |
| Outcome: | The proposed approach generates a lexicon that performs as well as the most advanced lexical induction method. |
Improving Sentence Boundary Detection for Spoken Language Transcripts (2020.lrec-1)
Copied to clipboard
| Challenge: | Using data expansion and transfer learning, we find that data expansion does not always improve results. |
| Approach: | They propose to divide spoken language into sentence-like units using Topological Fields model . they also propose to use data from the same domain to test different ML architectures . |
| Outcome: | The proposed model improves the detection of boundary detection in spoken dialogues compared to a sequence tagging approach. |
Detecting Derogatory Compounds – An Unsupervised Approach (N19-1)
Copied to clipboard
| Challenge: | Derogatory compounds are more difficult to detect than derogatory unigrams since they are sparsely represented in general-purpose lexical resources. |
| Approach: | They propose an unsupervised classification approach that incorporates linguistic properties of compounds. |
| Outcome: | The proposed method is compared with existing methods for extracting derogatory unigrams . the proposed method uses a distributional representation to incorporate linguistic properties of compounds . |
Doctor Who? Framing Through Names and Titles in German (2020.lrec-1)
Copied to clipboard
| Challenge: | Entity framing is the selection of aspects of an entity to promote a particular viewpoint towards that entity. |
| Approach: | They investigate entity framing of political figures through the use of names and titles in German online discourse. |
| Outcome: | The proposed method improves existing studies on German political discourse . it shows that the formality of naming correlates positively with stance in the tweets . |
Enhancing a Lexicon of Polarity Shifters through the Supervised Classification of Shifting Directions (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing polarity shifter lexica only specify when a word can cause shifting, but do not specify when this is limited to a single shifting direction. |
| Approach: | They propose a classifier that determines the shifting direction of polarity shifters by using resource-driven features and data-driven feature. |
| Outcome: | The proposed classifier enhances the largest available polarity shifter lexicon. |
Detection of Abusive Language: the Problem of Biased Datasets (N19-1)
Copied to clipboard
| Challenge: | Recent studies have reported high classification performance on datasets with difficult cases of abusive language. |
| Approach: | They examine the impact of data bias on abusive language detection by focusing on specific microposts rather than random sampling. |
| Outcome: | The proposed method is more accurate and more accurate than random sampling. |