Papers by Alexandra Olteanu
“One-Size-Fits-All”? Examining Expectations around What Constitute “Fair” or “Good” NLG System Behaviors (2024.naacl-long)
Copied to clipboard
| Challenge: | Natural language generation models are used for many downstream applications involving interpersonal communication, such as text completion, "smart" reply suggestions, and chatbot assistants. |
| Approach: | They conduct five case studies that perturb identity-related language features in NLG inputs to examine their assumptions about fairness. |
| Outcome: | The findings highlight open challenges around what constitutes “fair” or “good” NLG system behaviors. |
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have focused on how text generation systems can lead to harmful outcomes such as over-reliance, emotional dependence, dehumanization, deception, or even physical harm. |
| Approach: | They propose to use an inventory of interventions to help identify possible interventions and provide a conceptual framework to help characterize the landscape of possible interventions. |
| Outcome: | The proposed interventions are based on an inventory of interventions grounded in prior literature and a crowdsourcing study where participants edited system outputs to make them less human-like. |
Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications (2022.naacl-main)
Copied to clipboard
| Challenge: | Evaluating natural language generation systems is difficult, as there are many ways to express similar things in text. |
| Approach: | They combine interviews with NLG practitioners to examine ethical considerations and their implications for NLG evaluation. |
| Outcome: | The findings of the study surface goals, community practices, assumptions, and constraints that shape NLG evaluations, and examine their implications and how they embody ethical considerations. |
FairPrism: Evaluating Fairness-Related Harms in Text Generation (2023.acl-long)
Copied to clipboard
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, Hanna Wallach
| Challenge: | FairPrism dataset provides a framework for measuring and mitigating fairness-related harms caused by AI text generation systems. |
| Approach: | They propose a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality. |
| Outcome: | FairPrism is a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering harms relating to gender and sexuality. |
The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources (2023.acl-long)
Copied to clipboard
Akshatha Arodi, Martin Pömsl, Kaheer Suleman, Adam Trischler, Alexandra Olteanu, Jackie Chi Kit Cheung
| Challenge: | Existing models that make inferences using information from multiple sources are largely understudied . |
| Approach: | They propose a test suite of coreference resolution subtasks that require reasoning over multiple facts and introduce subtask where knowledge is present only at inference time using fictional knowledge. |
| Outcome: | The proposed subtasks differ in terms of which knowledge sources contain the relevant facts and where knowledge is present only at inference time using fictional knowledge. |
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets (2021.acl-long)
Copied to clipboard
| Challenge: | Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measurements of harms. |
| Approach: | They apply a measurement modeling lens to inventory pitfalls that threaten benchmarks' validity as measurement models for stereotyping. |
| Outcome: | The proposed benchmarks lack clarity and assumptions that affect how they conceptualize and operationalize stereotyping. |
Responsible AI Considerations in Text Summarization Research: A Review of Current Practices (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a recent study examines research and reporting practices for text summarization tasks . text summaries are often overlooked by the responsible AI community . |
| Approach: | They examine research and reporting practices in the context of text summarization . they find that relatively few papers engage with possible stakeholders . |
| Outcome: | The findings highlight current research practices and provide recommendations on research directions. |
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective (2024.findings-acl)
Copied to clipboard
| Challenge: | a recent study shows that evaluations of CR models on multiple datasets conflate different factors concerning what is being measured. |
| Approach: | They propose to view evaluations through the lens of measurement modeling . they show that evaluations risk conflating different factors concerning what is being measured . |
| Outcome: | The evaluations on seven datasets show that models that reflect coreference generalization are often correlated with differences in how coreference is defined and operationalized. |
ADEPT: An Adjective-Dependent Plausibility Task (2021.acl-long)
Copied to clipboard
| Challenge: | ADEPT is a large-scale semantic plausibility task that requires a significant degree of world knowledge and common-sense reasoning. |
| Approach: | They propose a large-scale semantic plausibility task that pairs 16 thousand sentences with slightly modified versions obtained by adding an adjective to a noun. |
| Outcome: | The proposed task is easier for humans (85% accuracy), but more difficult for transformer-based models (71% accuracy). |
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems (2025.findings-acl)
Copied to clipboard
Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, Hanna Wallach
| Challenge: | Existing tools for measuring representational harms caused by large language model systems are not useful for practitioners. |
| Approach: | They examine the extent to which public instruments are used to measure representational harms caused by large language model-based systems. |
| Outcome: | The proposed instruments do not meet the needs of practitioners evaluating large language model-based systems. |