Papers by Sharon Levy
Towards Understanding Gender-Seniority Compound Bias in Natural Language Generation (2022.lrec-1)
Copied to clipboard
Samhita Honnavalli, Aesha Parekh, Lily Ou, Sophie Groenwold, Sharon Levy, Vicente Ordonez, William Yang Wang
| Challenge: | Existing studies have not investigated how gender biases in natural language processing (NLP) are compounded with other societal biase. |
| Approach: | They propose a framework for probing compound bias by examining seniority in pre-trained neural generation models. |
| Outcome: | The proposed framework amplifies bias by considering women as junior and men as senior more often than ground truth in both domains. |
Characterizing Selective Refusal Bias in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | a recent study shows that safety guardrails in large language models can inadvertently introduce or reflect new biases as they may refuse to generate harmful content targeting some demographic groups and not others. |
| Approach: | They examine the selective refusal bias in large language models by examining demographics and responses. |
| Outcome: | The proposed model fails to defend against an indirect attack on previously refused groups in 89% of the trials. |
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies have shown that relying on LLMs as information providers may hurt student learning. |
| Approach: | They introduce and apply two bias score metrics to evaluate LLMs for bias in the personalized educational setting, specifically on the models’ roles as “teachers.” |
| Outcome: | The proposed models harm student learning by perpetuating harmful stereotypes and reversing them. |
Evaluating Biases in Context-Dependent Sexual and Reproductive Health Questions (2024.findings-emnlp)
Copied to clipboard
| Challenge: | With the rise in accessibility of chat-based large language models, the public increasingly uses them as question-answering systems for personalized answers. |
| Approach: | They curate a dataset of sexual and reproductive healthcare questions dependent on age, sex, and location attributes and compare their outputs with and without demographic context to determine answer alignment . |
| Outcome: | The results show that young adult female users are favored in the model answers to underspecified questions in the healthcare domain. |
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models acquire beliefs about gender from training data and can therefore generate text with stereotypical gender attitudes. |
| Approach: | They use a decision-making lens to examine gender equity within large language models . they explore relationships through typical and gender-neutral names . |
| Outcome: | The proposed model generation and classification models exhibit stereotypical gender biases . the proposed model generates gender-neutral names, with and without safety enhancements, and egalitarian versus traditional scenarios across topics. |
Comparing Biases and the Impact of Multilingual Training across Multiple Languages (2023.emnlp-main)
Copied to clipboard
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, Dan Roth
| Challenge: | Currently, studies on bias and fairness in natural language processing focus on a single language and/or across few attributes (e.g. gender, race). However, biases can manifest differently across languages for individual attributes. |
| Approach: | They adapt existing sentiment bias templates in English to Italian, Chinese, Hebrew, and Spanish for race, religion, nationality, and gender. |
| Outcome: | The proposed model favors groups that are dominant in each language's culture, indicating bias amplification, after multilingual finetuning. |
ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models do not provide robustness evaluations for large language models, but we find that they are inconsistent in performance. |
| Approach: | They propose to use semantically aligned augmentation, target bootstrapping, and adversarial knowledge injection to generate a test suite of prompts covering diverse robustness settings. |
| Outcome: | The proposed system generates a set of prompts covering diverse settings covering semantic equivalence, related scenarios, and adversarial. |
Making FETCH! Happen: Finding Emergent Dog Whistles Through Common Habitats (2025.acl-long)
Copied to clipboard
| Challenge: | Dog whistles are coded expressions with dual meanings that slip by content moderation filters . a new study finds that state-of-the-art systems fail to identify novel dog whistles . |
| Approach: | They propose a task to find novel dog whistles in massive social media corpora . they use a strong baseline system that combines vector databases and Large Language Models to identify new dog whistle. |
| Outcome: | The proposed system fails to identify dog whistles across three social media cases . it combines vector databases and Large Language Models to efficiently and effectively identify new dog whistle expressions. |
Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts (2024.naacl-short)
Copied to clipboard
| Challenge: | With growth in the popularity of text-to-image models has come interest in assessing their multilingual capabilities, including multilingual accessibility. |
| Approach: | They propose to correct translation errors in a concept list translated to seven languages and compare the outputs of the benchmark to those conditioned on the old. |
| Outcome: | The proposed benchmark contains translation errors in Spanish, Japanese, and Chinese. |
Addressing Issues of Cross-Linguality in Open-Retrieval Question Answering Systems For Emergent Domains (2023.eacl-demo)
Copied to clipboard
| Challenge: | a lack of cross-lingual training data in emergent domains makes it difficult to train on emerging domains. |
| Approach: | They propose a cross-lingual open-retrieval question answering system for COVID-19 . their system adopts a corpus of scientific articles to ensure that retrieved documents are reliable. |
| Outcome: | The proposed system outperforms BM25 baselines in cross-lingual settings. |
Mitigating Covertly Unsafe Text within Natural Language Systems (2022.findings-emnlp)
Copied to clipboard
Alex Mei, Anisha Kabir, Sharon Levy, Melanie Subbiah, Emily Allaway, John Judge, Desmond Patton, Bruce Bimber, Kathleen McKeown, William Yang Wang
| Challenge: | Existing studies on text safety have focused on overtly unsafe, covertly, or indirectly unsafe statements. |
| Approach: | They propose a method to identify physical harm-causing statements as overtly, covertly or indirectly unsafe and a solution to mitigate the generation of such statements. |
| Outcome: | The proposed methods identify the type of unsafe language that can cause physical harm and identify mitigation strategies to inspire future researchers to tackle this challenging problem. |
Modeling Disclosive Transparency in NLP Application Descriptions (2021.emnlp-main)
Copied to clipboard
| Challenge: | Broader disclosive transparency is difficult to define and quantify, authors say . previous work has demonstrated trade-offs and negative consequences to disclosing transparency . |
| Approach: | They propose to use neural language model-based probabilistic metrics to model disclosive transparency . they demonstrate that they correlate with user and expert opinions of system transparency a valid objective proxy . |
| Outcome: | The proposed metrics correlate with user and expert opinions of system transparency, making them a valid objective proxy. |
HybriDialogue: An Information-Seeking Dialogue Dataset Grounded on Tabular and Textual Data (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets focused on multiturn dialogue systems focus on text or table information. |
| Approach: | They propose a dataset that consists of crowdsourced conversations grounded on Wikipedia text and tables. |
| Outcome: | The proposed dataset shows that there is still ample opportunity for improvement in the current state of dialogue systems. |
Foveate, Attribute, and Rationalize: Towards Physically Safe and Trustworthy AI (2023.findings-acl)
Copied to clipboard
| Challenge: | Covertly unsafe text is an area of particular interest as it is difficult to detect as harmful . previous work focused on explicit violent text and typically expressed through violent keywords. |
| Approach: | They propose a framework that leverages external knowledge for trustworthy rationale generation in the context of safety. |
| Outcome: | The proposed framework improves safety classification accuracy by 5.9% on the SafeText dataset, and shows that it is more accurate than previous frameworks. |
Investigating African-American Vernacular English in Transformer-Based Text Generation (2020.emnlp-main)
Copied to clipboard
Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, William Yang Wang
| Challenge: | Recent work in Natural Language Generation (NLG) uses a Transformer-based language model to generate high-quality, coherent text when prompted by arbitrary input. |
| Approach: | They evaluate the performance of a Transformer-based model that generates high-quality, coherent text when prompted by arbitrary input. |
| Outcome: | The proposed model improves on AAVE and SAE text with pretrained sentiment classifiers. |
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to eliminate implicit biases in LLMs do not eradicate underlying behavioral bias. |
| Approach: | They propose a framework that uses logic grid puzzles to probe the influence of social stereotypes on logical reasoning and decision making in LLMs. |
| Outcome: | The proposed framework systematically probes the influence of social stereotypes on logical reasoning and decision making in LLMs. |
The Power of Summary-Source Alignments (2024.findings-acl)
Copied to clipboard
Ori Ernst, Ori Shapira, Aviv Slobodkin, Sharon Adar, Mohit Bansal, Jacob Goldberger, Ran Levy, Ido Dagan
| Challenge: | Multi-document summarization (MDS) is a challenging task, often decomposed to subtasks of salience and redundancy detection, followed by text generation. |
| Approach: | They propose to extend the summary-source alignment framework by applying it at the more fine-grained proposition span level and annotating alignment manually in a multi-document setup. |
| Outcome: | The proposed framework can yield several datasets for at least six different tasks. |
Open-Domain Question-Answering for COVID-19 and Other Emergent Domains (2021.emnlp-demo)
Copied to clipboard
| Challenge: | a system for open-domain question-answering is developed for COVID-19 . small data size allows system to retrieve answers from large corpus of scientific papers . |
| Approach: | They propose an open-domain question-answering system that can retrieve answers from large corpus of COVID-19 papers. |
| Outcome: | The proposed open-domain question-answering system can retrieve answers from large corpus of COVID-19 scientific papers. |
Investigating Memorization of Conspiracy Theories in Text Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies examine conspiracy theories in social media, but they have not evaluated their presence in generative language models. |
| Approach: | They examine the ability of generative language models to generate conspiracy theory text . they highlight the difficulties of this task and discuss the drawbacks . |
| Outcome: | The proposed model can generate conspiracy theories without access to training data. |
Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection (2020.lrec-1)
Copied to clipboard
| Challenge: | Prior fake news datasets lack multimodal text and image data, metadata, comment data, and fine-grained classification at the scale and breadth of their datasets. |
| Approach: | They propose to use a multimodal dataset to build a machine learning classification model that uses text and image data to classify fake news. |
| Outcome: | The proposed model is based on a multimodal dataset consisting of over 1 million samples from multiple categories of fake news. |
SafeText: A Benchmark for Exploring Physical Safety in Language Models (2022.emnlp-main)
Copied to clipboard
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang
| Challenge: | Existing models that generate unsafe text are susceptible to the dangers of unsafe text generation and are deemed unsafe. |
| Approach: | They use a dataset to empirically study commonsense physical safety across various models for text generation and reasoning tasks. |
| Outcome: | The proposed model can generate unsafe text and reject it, but the different harms that can occur do not receive equal attention, which may consequently downplay certain harms. |