Papers by Shachi Dave

5 papers
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets on social stereotypes are limited in size and coverage . existing datasets are restricted to stereotypes prevalent in the Western society .
Approach: They propose a broad-coverage stereotype dataset using generative models and a globally diverse rater pool to validate the prevalence of stereotypes in society.
Outcome: The dataset validates the prevalence of stereotypes in society across 8 geo-political regions across 6 continents and states within the US and India.
Parameter-Efficient Finetuning for Robust Continual Multilingual Learning (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to Continual Multilingual Learning (CML) are based on updating models using new data in stages.
Approach: They propose a parameter-efficient finetuning strategy to increase the number of languages on which the model improves after an update while reducing the magnitude of loss for the remaining languages.
Outcome: The proposed model improves on the languages included in the latest update while reducing the loss of performance on the remaining languages.
Re-contextualizing Fairness in NLP: The Case of India (2022.aacl-main)

Copied to clipboard

Challenge: Recent research has revealed undesirable biases in NLP data and models . however, these efforts focus of social disparities in the West and are not directly portable to other geo-cultural contexts.
Approach: They propose a framework to re-contextualize NLP fairness research for the Indian context . they build resources for fairness evaluation in the Indian and delve deeper into social stereotypes for Region and Religion .
Outcome: The proposed framework can be generalized to other geo-cultural contexts.
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes.
Approach: They propose to use a dataset to evaluate nationality-based stereotypes in T2I models across 135 nationalities to assess offensive stereotypes.
Outcome: The proposed dataset enables evaluation of known nationality-based stereotypes across 135 nationalities.
Bootstrapping Multilingual Semantic Parsers using Large Language Models (2023.eacl-main)

Copied to clipboard

Challenge: Despite cross-lingual generalization, translation models require significant amounts of labeled data for many low-resource languages . brittle translation services may be due to domain mismatch between input text and general-purpose text .
Approach: They propose to use large language models to translate English datasets into several languages via few-shot prompting.
Outcome: The proposed method outperforms a strong translation-train baseline on 41 out of 50 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations