Papers by Ioana Baldini

7 papers
Domain Generalizable AI Guardrails with Augmented Policy Training (2026.acl-long)

Copied to clipboard

Challenge: Current guardrails overfit the training policies, preventing adaptation to new domains and policies.
Approach: They propose a training recipe that uses a suite of policy perturbation strategies to reduce overfitting and increase generalization to guardrails.
Outcome: The proposed training recipe reduces overfitting and increases generalization on unseen policies and achieves comparable or better performance than existing 8B guardrails on unsen policies.
Biomedical Interpretable Entity Representations (2021.findings-acl)

Copied to clipboard

Challenge: Existing work on general interpretable representation learning does not transfer to biomedicine . pre-trained models induce dense entity representations but are not immediately interpretable.
Approach: They propose a method that exploits BIER's final sparse and intermediate dense representations to facilitate model and entity type debugging.
Outcome: The proposed model performs well on biomedical tasks including disambiguation and label classification.
Biasly: An Expert-Annotated Dataset for Subtle Misogyny Detection and Mitigation (2024.findings-acl)

Copied to clipboard

Challenge: the Biasly dataset captures misogyny in movies in ways unique within the literature.
Approach: The Biasly dataset captures misogyny in North American film by combining annotations of movie subtitles with common NLP algorithms.
Outcome: The Biasly dataset captures misogyny expressions in North American film . it contains annotations of movie subtitles and text generation for rewrites .
DAMAGeR: Deploying Automatic and Manual Approaches to GenAI Red-teaming (2025.naacl-tutorial)

Copied to clipboard

Challenge: In this tutorial, we will review and apply current automatic and manual red-teaming techniques for GenAI models.
Approach: This tutorial will review automatic and manual red-teaming techniques for GenAI models .
Outcome: This tutorial will review and apply current automatic and manual red-teaming techniques for GenAI models.
Why Don’t Prompt-Based Fairness Metrics Correlate? (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to assess fairness using prompts have low correlations between fairness metrics.
Approach: They propose a method to enhance the correlation between fairness metrics by using pre-trained language models.
Outcome: The proposed method improves the correlation between fairness metrics by using pre-trained language models.
Your fairness may vary: Pretrained language model fairness in toxic text classification (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained, bidirectional language models have revolutionized natural language processing research . authors show that focusing on accuracy measures alone can lead to models with wide variation in fairness characteristics .
Approach: They propose to use two post-processing methods to improve model fairness without retraining . they use pretrained language models of varying sizes on two toxic text classification tasks .
Outcome: The proposed methods improve model fairness without retraining . the results show that the fairness variation is more than just accuracy .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations