Papers by Sunipa Dev
Representation Learning for Resource-Constrained Keyphrase Generation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | State-of-the-art keyphrase generation methods depend on large annotated datasets, limiting their performance in domains with limited annotation data. |
| Approach: | They propose a method that first identifies salient information using retrieval-based corpus-level statistics and then learns a task-specific intermediate representation based on a pre-trained language model. |
| Outcome: | The proposed method improves keyphrase generation and zero-shot domain adaptation on multiple keyphrase benchmarks. |
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context (2026.eacl-short)
Copied to clipboard
Aishwarya Verma, Laud Ammah, Olivia Nercy Ndlovu Lucas, Andrew Zaldivar, Vinodkumar Prabhakaran, Sunipa Dev
| Challenge: | Existing data collection approaches to generative AI are inadequate to assess its safety and utility. |
| Approach: | They propose a multilingual stereotype resource that uses socioculturally-situated, community-engaged methods to assess the region’s linguistic diversity and traditional orality. |
| Outcome: | The proposed method covers four sub-Saharan African countries that are severely underrepresented in NLP resources: Ghana, Kenya, Nigeria, and South Africa. |
MisgenderMender: A Community-Informed Approach to Interventions for Misgendering (2024.naacl-long)
Copied to clipboard
| Challenge: | Misgendering is the act of incorrectly addressing someone’s gender and is pervasive in everyday use platforms and technologies. |
| Approach: | They propose a task and evaluation dataset to assess the effectiveness of automated misgendering interventions for text-based misgending in the US. |
| Outcome: | The proposed dataset includes 3790 instances of social media content and LLM-generations about non-cisgender public figures, annotated for the presence of misgendering, with additional annotations for correcting misgending in LLM generated text. |
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models (2023.acl-long)
Copied to clipboard
Akshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, Sunipa Dev
| Challenge: | Existing datasets on social stereotypes are limited in size and coverage . existing datasets are restricted to stereotypes prevalent in the Western society . |
| Approach: | They propose a broad-coverage stereotype dataset using generative models and a globally diverse rater pool to validate the prevalence of stereotypes in society. |
| Outcome: | The dataset validates the prevalence of stereotypes in society across 8 geo-political regions across 6 continents and states within the US and India. |
MiTTenS: A Dataset for Evaluating Gender Mistranslation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on gender mistranslation in translation systems have highlighted the problem . a dataset of 26 languages is presented to measure the extent of such errors . |
| Approach: | They propose a dataset that measures the extent of gender mistranslation in translation systems . they use handcrafted passages that target known failure patterns and synthetically generated passages . |
| Outcome: | The proposed dataset covers 26 languages from a variety of language families and scripts, including several traditionally under-represented in digital resources. |
Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES) (2026.findings-acl)
Copied to clipboard
Guido Ivetta, Pietro Palombini, Sofía Martinelli, Marcos J Gomez, M Emilia Echeveste, Sunipa Dev, Vinodkumar Prabhakaran, Luciana Benotti
| Challenge: | a geo-cultural gap in NLP evaluation hinders evaluation of societal biases . authors propose a new method to collect stereotypes from large language models . |
| Approach: | They propose a new method that integrates sourcing and validation of existing data into a single workflow. |
| Outcome: | The proposed method improves LACES by integrating new stereotype entries and validation of existing data. |
Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent work analyzes, quantifies, and mitigates language model biases such as gender, race or religion-related stereotypes in static word embeddings and contextual representations. |
| Approach: | They explain the complexity of gender and language around it and examine how current representations perpetuate harms associated with binary gender. |
| Outcome: | The proposed model and dataset biases perpetuate harms associated with the treatment of gender as binary in English language technologies. |
MISGENDERED: Limits of Large Language Models in Understanding Pronouns (2023.acl-long)
Copied to clipboard
| Challenge: | excluding non-binary gender identities can perpetuate harm against non-bisexual individuals through exclusion and marginalization. |
| Approach: | They propose a framework for evaluating large language models’ ability to correctly use preferred pronouns. |
| Outcome: | The proposed framework evaluates language models' ability to correctly use preferred pronouns in English. |
On Measures of Biases and Harms in NLP (2022.findings-aacl)
Copied to clipboard
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, Kai-Wei Chang
| Challenge: | Recent studies show that natural language processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality. |
| Approach: | They propose a framework for harms and questions to help practitioners understand biases . they propose measurable measures to detect and mitigate biased groups . |
| Outcome: | The proposed framework provides a framework for harms and questions for practitioners to answer to guide the development of bias measures. |
Amplifying Trans and Nonbinary Voices: A Community-Centred Harm Taxonomy for LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies on harms of language technology to transgender and nonbinary people focus on misgendering and stereotyping . |
| Approach: | They propose a taxonomy of harms for large language models and heuristics for evaluation to help identify harmful behavior in LLMs. |
| Outcome: | The proposed model-based approach combines surveys and focus groups with community experts to identify harmful behavior in large language models. |
The Tail Wagging the Dog: Dataset Construction Biases of Social Bias Benchmarks (2023.acl-short)
Copied to clipboard
| Challenge: | omnipresence of large pre-trained language models has fueled concerns regarding systematic biases carried over from underlying data into the applications they are used in. |
| Approach: | They propose to compare social biases with non-social biase masked by alternate constructions that maintain the essence of their social bias. |
| Outcome: | The proposed benchmarks underestimate or overestimate the social bias in a given model. |
A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent years have seen unprecedented gains in generative AI models' capabilities across modalitieslanguage, image, audio, and video domains across the globe. |
| Approach: | They propose a framework to operationalize stereotypes in generative AI evaluations using social psychological research and NLP data. |
| Outcome: | The proposed framework identifies key components of stereotypes that are crucial in AI evaluation, including the target group, associated attribute, relationship characteristics, perceiving group, and context. |
Scaling Cultural Resources for Improving Generative Models (2026.findings-eacl)
Copied to clipboard
Hayk Stepanyan, Aishwarya Verma, Andrew Zaldivar, Rutledge Chin Feman, Erin MacMurray van Liemt, Charu Kalia, Vinodkumar Prabhakaran, Sunipa Dev
| Challenge: | generative models have been known to have reduced performance in different global cultural contexts and languages. |
| Approach: | They construct a pipeline to collect and contribute culturally salient, multilingual data . they argue such data can assess the state of the global applicability of generative AI models . |
| Outcome: | The proposed pipeline can assess the state of the global applicability of our models and improve upon cross-cultural gaps. |
Socially Aware Bias Measurements for Hindi Language Representations (2022.naacl-main)
Copied to clipboard
| Challenge: | Language representations are an efficient tool used across NLP, but they are strife with encoded societal biases. |
| Approach: | They investigate the encoded biases in Hindi language representations based on cultural and historical contexts . they emphasize the necessity of social-awareness along with linguistic and grammatical artefacts when modeling language representation . |
| Outcome: | The proposed model reflects the cultural and cultural diversity of the region in which it is used . the model is based on the language and culture of the language being used based upon the study . |
Re-contextualizing Fairness in NLP: The Case of India (2022.aacl-main)
Copied to clipboard
| Challenge: | Recent research has revealed undesirable biases in NLP data and models . however, these efforts focus of social disparities in the West and are not directly portable to other geo-cultural contexts. |
| Approach: | They propose a framework to re-contextualize NLP fairness research for the Indian context . they build resources for fairness evaluation in the Indian and delve deeper into social stereotypes for Region and Religion . |
| Outcome: | The proposed framework can be generalized to other geo-cultural contexts. |
OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to mitigate stereotypical biases by linear projection are too aggressive . existing methods remove bias, but they also erase valuable information from word embeddings . |
| Approach: | They propose a bias-mitigating method that disentangles biased associations between concepts instead of removing concepts wholesale. |
| Outcome: | The proposed method disentangles biased associations between concepts rather than eliminating concepts wholesale. |
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation (2024.acl-long)
Copied to clipboard
Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan Reddy, Sunipa Dev
| Challenge: | Existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes. |
| Approach: | They propose to use a dataset to evaluate nationality-based stereotypes in T2I models across 135 nationalities to assess offensive stereotypes. |
| Outcome: | The proposed dataset enables evaluation of known nationality-based stereotypes across 135 nationalities. |
Towards Geo-Culturally Grounded LLM Generations (2025.acl-short)
Copied to clipboard
| Challenge: | Contemporary large language models (LLMs) are pretrained on huge corpora of natural language text and fine-tuned using human feedback to improve their quality. |
| Approach: | They compare the performance of standard LLMs, LLM augmented with retrievals from a bespoke knowledge base and LLM with retrieval from . a web search on multiple cultural awareness benchmarks. |
| Outcome: | The retrieval augmented generation and search grounding techniques improve LLMs' ability to display familiarity with various national cultures on cultural awareness benchmarks. |
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies . |
| Approach: | They will examine how different socio-cultural perspectives influence what is taken as ground truth by models. |
| Outcome: | This tutorial examines how different socio-cultural perspectives influence representations of global concepts. |