Papers by Ioana Baldini
Domain Generalizable AI Guardrails with Augmented Policy Training (2026.acl-long)
Copied to clipboard
| Challenge: | Current guardrails overfit the training policies, preventing adaptation to new domains and policies. |
| Approach: | They propose a training recipe that uses a suite of policy perturbation strategies to reduce overfitting and increase generalization to guardrails. |
| Outcome: | The proposed training recipe reduces overfitting and increases generalization on unseen policies and achieves comparable or better performance than existing 8B guardrails on unsen policies. |
Biomedical Interpretable Entity Representations (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing work on general interpretable representation learning does not transfer to biomedicine . pre-trained models induce dense entity representations but are not immediately interpretable. |
| Approach: | They propose a method that exploits BIER's final sparse and intermediate dense representations to facilitate model and entity type debugging. |
| Outcome: | The proposed model performs well on biomedical tasks including disambiguation and label classification. |
Biasly: An Expert-Annotated Dataset for Subtle Misogyny Detection and Mitigation (2024.findings-acl)
Copied to clipboard
Brooklyn Sheppard, Anna Richter, Allison Cohen, Elizabeth Smith, Tamara Kneese, Carolyne Pelletier, Ioana Baldini, Yue Dong
| Challenge: | the Biasly dataset captures misogyny in movies in ways unique within the literature. |
| Approach: | The Biasly dataset captures misogyny in North American film by combining annotations of movie subtitles with common NLP algorithms. |
| Outcome: | The Biasly dataset captures misogyny expressions in North American film . it contains annotations of movie subtitles and text generation for rewrites . |
DAMAGeR: Deploying Automatic and Manual Approaches to GenAI Red-teaming (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | In this tutorial, we will review and apply current automatic and manual red-teaming techniques for GenAI models. |
| Approach: | This tutorial will review automatic and manual red-teaming techniques for GenAI models . |
| Outcome: | This tutorial will review and apply current automatic and manual red-teaming techniques for GenAI models. |
Why Don’t Prompt-Based Fairness Metrics Correlate? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods to assess fairness using prompts have low correlations between fairness metrics. |
| Approach: | They propose a method to enhance the correlation between fairness metrics by using pre-trained language models. |
| Outcome: | The proposed method improves the correlation between fairness metrics by using pre-trained language models. |
SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models (2025.naacl-long)
Copied to clipboard
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, null Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |
Your fairness may vary: Pretrained language model fairness in toxic text classification (2022.findings-acl)
Copied to clipboard
| Challenge: | Pre-trained, bidirectional language models have revolutionized natural language processing research . authors show that focusing on accuracy measures alone can lead to models with wide variation in fairness characteristics . |
| Approach: | They propose to use two post-processing methods to improve model fairness without retraining . they use pretrained language models of varying sizes on two toxic text classification tasks . |
| Outcome: | The proposed methods improve model fairness without retraining . the results show that the fairness variation is more than just accuracy . |