Papers by Hannah Kirk
A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning (2022.aacl-main)
Copied to clipboard
| Challenge: | Large-scale, pretrained vision-language models are growing in popularity due to impressive performance on downstream tasks with minimal finetuning. |
| Approach: | They propose to apply ranking metrics to image-text representations to investigate bias measures and debiasing methods to reduce various bias measures. |
| Outcome: | The proposed model reduces bias measures with minimal degradation to image-text representations. |
Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing models for detecting hate expressed with emojis have weaknesses when used for sensitive applications such as content moderation. |
| Approach: | They propose a test suite of 3,930 short-form statements that evaluates hateful language expressed with emoji. |
| Outcome: | The proposed model performs better on emoji-based hate while maintaining strong performance on text-only hate. |
Handling and Presenting Harmful Text in NLP Research (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Text data can pose a risk of harm, but the risks remain unresolved in the NLP community. |
| Approach: | They propose an analytical framework categorising harms on three axes: harm type, whether harm sought as a feature of research design, whether harmful content is encountered when working on unrelated problems, and who it affects . |
| Outcome: | The proposed framework categorises harms on three axes: harm type, whether harm sought as feature of research design, and whether harmful content is encountered when working on unrelated problems. |
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values (2023.emnlp-main)
Copied to clipboard
| Challenge: | Incorporating human feedback into Large Language Models is a welcome development, but it introduces new biases and challenges. |
| Approach: | They propose to survey 95 articles that use human feedback to steer, guide or tailor the behaviours of large language models. |
| Outcome: | The proposed approaches are based on 95 articles primarily from the ACL and arXiv repositories and highlight five unresolved conceptual and practical challenges. |
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are now being used by millions of people across the world. |
| Approach: | They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way. |
| Outcome: | The proposed test suite identifies eXaggerated Safety behaviours in a systematic way. |