Papers by Dominique Fohr
Transferring Knowledge via Neighborhood-Aware Optimal Transport for Low-Resource Hate Speech Detection (2022.aacl-main)
Copied to clipboard
| Challenge: | Existing approaches to detect hate speech are expensive and time-consuming . a new approach allows for flexible learning of neighborhood information . |
| Approach: | They propose a method that allows flexible modeling of neighbors retrieved from a resource-rich corpus to learn the amount of transfer. |
| Outcome: | The proposed training strategy improves on low-resource hate speech corpora over baselines. |
Domain Classification-based Source-specific Term Penalization for Domain Adaptation in Hate-speech Detection (2022.coling-1)
Copied to clipboard
| Challenge: | Existing approaches for hate-speech detection exhibit poor performance in out-of-domain settings due to overemphasizing source-specific information that negatively impacts its domain invariance. |
| Approach: | They propose a domain adaptation approach that automatically extracts and penalizes source-specific terms using a classifier. |
| Outcome: | The proposed approach improves cross-domain evaluation on indomain held-out instances while preserving high performance on out-of-domain settings. |
Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate Online (2022.lrec-1)
Copied to clipboard
Dana Ruiter, Liane Reiners, Ashwin Geet D’Sa, Thomas Kleinbauer, Dominique Fohr, Irina Illina, Dietrich Klakow, Christian Schemer, Angeliki Monnier
| Challenge: | HS-related corpora over-simplify the phenomenon of hate by labelling user content with binary classes, e.g., hate/neutral . this ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corporales. |
| Approach: | They present a corpus of 9k German and french user comments from migration-related news articles. |
| Outcome: | The proposed corpus is annotated with 23 features that become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate. |
Identification of Multiword Expressions in Tweets for Hate Speech Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | Multiword expression (MWE) identification in tweets is a complex task due to the complex linguistic nature of MWEs combined with the non-standard language use in social networks. |
| Approach: | They propose a new architecture for incorporating multiword expression features into tweets to improve their accuracy. |
| Outcome: | The proposed system outperforms existing systems on the hate speech detection task on English Twitter. |
Dynamically Refined Regularization for Improving Cross-corpora Hate Speech Detection (2022.findings-acl)
Copied to clipboard
| Challenge: | Hate speech classifiers exhibit performance degradation when evaluated on datasets different from the source. |
| Approach: | They propose to automatically identify and reduce spurious correlations using attribution methods with dynamic refinement of the list of terms that need to be regularized during training. |
| Outcome: | The proposed method improves performance across corpora and on different datasets. |