Papers by Paul Rottger
Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing models for detecting hate expressed with emojis have weaknesses when used for sensitive applications such as content moderation. |
| Approach: | They propose a test suite of 3,930 short-form statements that evaluates hateful language expressed with emoji. |
| Outcome: | The proposed model performs better on emoji-based hate while maintaining strong performance on text-only hate. |
Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks (2022.naacl-main)
Copied to clipboard
| Challenge: | Labelled data is the foundation of most natural language processing tasks, but there are valid beliefs about what the correct data labels should be. |
| Approach: | They propose two contrasting paradigms for data annotation that encourage annotator subjectivity . they propose a descriptive paradigm that allows for the surveying and modelling of different beliefs . |
| Outcome: | The proposed paradigms encourage annotator subjectivity, while the prescriptive paradigm discourages it. |
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values (2023.emnlp-main)
Copied to clipboard
| Challenge: | Incorporating human feedback into Large Language Models is a welcome development, but it introduces new biases and challenges. |
| Approach: | They propose to survey 95 articles that use human feedback to steer, guide or tailor the behaviours of large language models. |
| Outcome: | The proposed approaches are based on 95 articles primarily from the ACL and arXiv repositories and highlight five unresolved conceptual and practical challenges. |