Papers by A. Bergman
SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems (2022.acl-long)
Copied to clipboard
| Challenge: | Several studies discuss the potential harms and benefits of large language models (LLMs) large neural models can replicate and even amplify negative, stereotypical, and derogatory associations in the data. |
| Approach: | They propose to use a first aid kit to assess the safety of conversational AI in various settings . they propose several future directions and discuss ethical considerations . |
| Outcome: | The proposed tools can provide estimates of the relative safety of systems in various settings, but they still have several shortcomings. |
STAR: SocioTechnical Approach to Red Teaming Language Models (2024.emnlp-main)
Copied to clipboard
Laura Weidinger, John Mellor, Bernat Pegueroles, Nahema Marchal, Ravin Kumar, Kristian Lum, Canfer Akbulut, Mark Diaz, A. Bergman, Mikel Rodriguez, Verena Rieser, William Isaac
| Challenge: | STAR is a sociotechnical framework that improves on current best practices for red teaming safety of large language models. |
| Approach: | They propose a sociotechnical framework that improves on current best practices for red teaming safety of large language models. |
| Outcome: | The proposed framework improves on current best practices for red teaming safety of large language models. |
Towards Responsible Natural Language Annotation for the Varieties of Arabic (2022.findings-acl)
Copied to clipboard
| Challenge: | In NLP, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance. |
| Approach: | They propose a playbook for responsible dataset creation for polyglossic, multidialectal languages . they focus on Arabic annotation of social media content as an example . |
| Outcome: | The proposed model is based on Arabic annotation of social media content. |