Papers by Hannah Kirk

5 papers
A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning (2022.aacl-main)

Copied to clipboard

Challenge: Large-scale, pretrained vision-language models are growing in popularity due to impressive performance on downstream tasks with minimal finetuning.
Approach: They propose to apply ranking metrics to image-text representations to investigate bias measures and debiasing methods to reduce various bias measures.
Outcome: The proposed model reduces bias measures with minimal degradation to image-text representations.
Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate (2022.naacl-main)

Copied to clipboard

Challenge: Existing models for detecting hate expressed with emojis have weaknesses when used for sensitive applications such as content moderation.
Approach: They propose a test suite of 3,930 short-form statements that evaluates hateful language expressed with emoji.
Outcome: The proposed model performs better on emoji-based hate while maintaining strong performance on text-only hate.
Handling and Presenting Harmful Text in NLP Research (2022.findings-emnlp)

Copied to clipboard

Challenge: Text data can pose a risk of harm, but the risks remain unresolved in the NLP community.
Approach: They propose an analytical framework categorising harms on three axes: harm type, whether harm sought as a feature of research design, whether harmful content is encountered when working on unrelated problems, and who it affects .
Outcome: The proposed framework categorises harms on three axes: harm type, whether harm sought as feature of research design, and whether harmful content is encountered when working on unrelated problems.
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values (2023.emnlp-main)

Copied to clipboard

Challenge: Incorporating human feedback into Large Language Models is a welcome development, but it introduces new biases and challenges.
Approach: They propose to survey 95 articles that use human feedback to steer, guide or tailor the behaviours of large language models.
Outcome: The proposed approaches are based on 95 articles primarily from the ACL and arXiv repositories and highlight five unresolved conceptual and practical challenges.
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are now being used by millions of people across the world.
Approach: They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way.
Outcome: The proposed test suite identifies eXaggerated Safety behaviours in a systematic way.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations