Challenge: Existing benchmark datasets focus on English language and the Western context, leaving a void for a reliable dataset that encapsulates India’s unique socio-cultural nuances.
Approach: They propose to use CrowS-Pairs to create a benchmark dataset that captures and evaluates social biases in Large Language Models (LLMs).
Outcome: The proposed dataset is available in English and Hindi and leverages LLMs ChatGPT and InstructGPT to augment the existing dataset with diverse societal biases and stereotypes prevalent in India.

Similar Papers

FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on fairness of LLMs are largely Western-focused, making them inadequate for culturally diverse countries such as India.
Approach: They propose a benchmark to evaluate fairness of LLMs across 85 identity groups . they consult domain experts to curate over 1,800 socio-cultural topics .
Outcome: The benchmark evaluates LLMs across 85 identities across 85 castes, religions, regions, and tribes.
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models use cultural biases implicitly, causing harm . identifying and quantifying learnt biase enables us to measure progress .
Approach: They propose a benchmark to measure social bias in pretrained language models . they use 1508 examples that cover stereotypes dealing with nine types of bias .
Outcome: The proposed benchmark focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups.
Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on human biases are heavily skewed towards Western and European languages . despite growing interest in language models, there are several shortcomings in the literature .
Approach: They scale the Word Embedding Association Test to 24 languages and add culturally relevant information for each language.
Outcome: The proposed language models can reflect and often amplify the effects of bias across linguistic, cultural, and societal borders.
BanStereoSet: A Dataset to Measure Stereotypical Social Biases in LLMs for Bangla (2025.findings-acl)

Copied to clipboard

Challenge: ***BanStereoSet*** is a dataset designed to evaluate stereotypical social biases in multilingual LLMs for the Bangla language.
Approach: They propose to localize the content from StereoSet, IndiBias, and kamruzzaman-etal's datasets to capture biases prevalent within the Bangla language.
Outcome: The proposed dataset consists of 1,194 sentences spanning 9 categories of bias: race, profession, gender, ageism, beauty, beauty in profession, region, caste, and religion.
TWBias: A Benchmark for Assessing Social Bias in Traditional Chinese Large Language Models through a Taiwan Cultural Lens (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models have shown remarkable capabilities in natural language processing, but concerns about social bias amplification remain.
Approach: They propose a social bias evaluation benchmark for Traditional Chinese LLMs that integrates chat templates and diverse prompts for comprehensive bias assessment.
Outcome: The proposed model incorporates chat templates and diverse prompts for comprehensive bias assessment focusing on Taiwan's cultural context and prioritizing gender and ethnicity bias evaluation.
Social Bias Probing: Fairness Benchmarking for Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluating social biases in language models have been limited to binary association tests on small datasets.
Approach: They propose a framework for probing language models for social biases by assessing disparate treatment . they use a large-scale benchmark to examine the diversity of identities and stereotypes .
Outcome: The proposed framework expands the analysis beyond the binary comparison of stereotypical versus anti-stereotypical identities to include a diverse range of identities and stereotypes.
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America (2025.emnlp-main)

Copied to clipboard

Challenge: a dataset of 46,499 sentences created in a professional development course captures intersectional biases across multiple demographic axes and school subjects.
Approach: They present a large-scale dataset of 46,499 sentences created in a professional development course . they show that the dataset contains more stereotypes unrecognized by current LLMs .
Outcome: The proposed dataset captures intersectional biases across multiple demographic axes and school subjects.
Socially Aware Bias Measurements for Hindi Language Representations (2022.naacl-main)

Copied to clipboard

Challenge: Language representations are an efficient tool used across NLP, but they are strife with encoded societal biases.
Approach: They investigate the encoded biases in Hindi language representations based on cultural and historical contexts . they emphasize the necessity of social-awareness along with linguistic and grammatical artefacts when modeling language representation .
Outcome: The proposed model reflects the cultural and cultural diversity of the region in which it is used . the model is based on the language and culture of the language being used based upon the study .
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English (2022.acl-long)

Copied to clipboard

Challenge: We introduce 1,679 sentence pairs in French that cover stereotypes in ten types of bias like gender and age.
Approach: They build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages and cultures.
Outcome: The proposed dataset allows for comparability across languages while characterizing biases that are specific to each country and language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations