Papers by Vinodkumar Prabhakaran

21 papers
Distinguishing Address vs. Reference Mentions of Personal Names in Text (2023.findings-acl)

Copied to clipboard

Challenge: Named entity recognition (NER) is a core task in the NLP community . but not much work has been done to distinguish between addressing and referring to entities .
Approach: They propose an automatic tagger that captures the address vs. reference distinction in English . they demonstrate how this distinction is important in NLP and computational social science applications .
Outcome: The proposed tagger performs at 85% accuracy in distinguishing between address and reference in English . many modern Indo-European languages do not have such vocative case markers .
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context (2026.eacl-short)

Copied to clipboard

Challenge: Existing data collection approaches to generative AI are inadequate to assess its safety and utility.
Approach: They propose a multilingual stereotype resource that uses socioculturally-situated, community-engaged methods to assess the region’s linguistic diversity and traditional orality.
Outcome: The proposed method covers four sub-Saharan African countries that are severely underrepresented in NLP resources: Ghana, Kenya, Nigeria, and South Africa.
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives (2024.naacl-long)

Copied to clipboard

Challenge: Recent work shows that ignoring rater subjectivity is problematic within specific tasks and for specific subgroups.
Approach: They propose a disagreement analysis framework to measure group association in perspectives among different rater subgroups.
Outcome: The proposed framework reveals specific rater groups that have significantly different perspectives than others on certain tasks and helps identify demographic axes that are crucial to consider in specific task contexts.
RtGender: A Corpus for Studying Differential Responses to Gender (L18-1)

Copied to clipboard

Challenge: Prior work on linguistic gender difference and communications about gender has focused on language about or portraying persons of a particular gender.
Approach: They present a multi-genre corpus of 25M comments from five socially and topically diverse sources tagged for the gender of the addressee and 30k annotations for sentiment and relevance of these responses.
Outcome: The proposed dataset shows that responses to women are more emotive and about the speaker as an individual (rather than about the content being responded to).
Social Biases in NLP Models as Barriers for Persons with Disabilities (2020.acl-main)

Copied to clipboard

Challenge: toxicity prediction and sentiment analysis models perpetuate undesirable social biases from the data on which they are trained.
Approach: They propose to use toxicity prediction and sentiment analysis to examine whether NLP models perpetuate undesirable biases towards mentions of disability.
Outcome: The proposed models contain undesirable biases towards mentions of disability in two English language models.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
Learning to Recognize Dialect Features (2021.naacl-main)

Copied to clipboard

Challenge: linguistics do not characterize dialects as simple categories, but as collections of correlated features.
Approach: They propose two multitask learning approaches based on pretrained transformers to detect dialect features in speech and text.
Outcome: The proposed models learn to recognize many features with high accuracy on 22 dialect features of Indian English.
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models (2023.acl-long)

Copied to clipboard

Challenge: Existing datasets on social stereotypes are limited in size and coverage . existing datasets are restricted to stereotypes prevalent in the Western society .
Approach: They propose a broad-coverage stereotype dataset using generative models and a globally diverse rater pool to validate the prevalence of stereotypes in society.
Outcome: The dataset validates the prevalence of stereotypes in society across 8 geo-political regions across 6 continents and states within the US and India.
Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES) (2026.findings-acl)

Copied to clipboard

Challenge: a geo-cultural gap in NLP evaluation hinders evaluation of societal biases . authors propose a new method to collect stereotypes from large language models .
Approach: They propose a new method that integrates sourcing and validation of existing data into a single workflow.
Outcome: The proposed method improves LACES by integrating new stereotype entries and validation of existing data.
Socially Responsible NLP (N18-6)

Copied to clipboard

Challenge: This tutorial will provide an overview of ethical research tools and ethical implications of language technologies.
Approach: This tutorial will provide an overview of ethical research and practical examples . it will discuss ethical tools to ensure data, algorithms, and models are socially responsible .
Outcome: This tutorial will provide an overview of ethical research tools and methods . it will discuss philosophical foundations of ethical work along with state of the art techniques .
D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and Evaluation (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies on annotator subjectivity focus on Western contexts and only document differences across age, gender, or racial groups.
Approach: They propose a large-scale cross-cultural dataset of parallel annotations for offensive language in over 4.5K English sentences annotated by a pool of more than 4k annotators from 21 countries.
Outcome: The proposed dataset captures annotators’ moral values along six moral foundations: care, equality, proportionality, authority, loyalty, and purity.
Perturbation Sensitivity Analysis to Detect Unintended Model Biases (D19-1)

Copied to clipboard

Challenge: Recent research shows that data-driven NLP models may inadvertently capture, reflect and sometimes amplify various social biases present in the language data they are trained on.
Approach: They propose a generic evaluation framework that detects unintended model biases related to named entities and requires no new annotations or corpora.
Outcome: The proposed framework detects unintended model biases related to named entities and requires no new annotations or corpora.
Author Commitment and Social Power: Automatic Belief Tagging to Infer the Social Context of Interactions (N18-1)

Copied to clipboard

Challenge: Social power is a difficult concept to define, but is often manifested in how we interact with one another.
Approach: They employ extra-propositional semantics extraction within NLP to study author commitment . they find that subordinates use significantly more instances of non-commitment than superiors .
Outcome: The proposed method shows that subordinates use significantly more instances of non-commitment than superiors, and that enriching lexical features with commitment labels captures important distinctions in social meanings.
BeSt: The Belief and Sentiment Corpus (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of propositional content is a set of cognitive attitudes of different agents towards a text . propositional attitudes are a cognitive attitude, including belief and sentiment, towards .
Approach: They propose a corpus which records cognitive state: who believes what, who has what sentiment . they use newswire and discussion forums in Chinese, English, and Spanish .
Outcome: The proposed corpus records who believes what (i.e., factuality) and who has what sentiment towards what.
A Comprehensive Framework to Operationalize Social Stereotypes for Responsible AI Evaluations (2025.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen unprecedented gains in generative AI models' capabilities across modalitieslanguage, image, audio, and video domains across the globe.
Approach: They propose a framework to operationalize stereotypes in generative AI evaluations using social psychological research and NLP data.
Outcome: The proposed framework identifies key components of stereotypes that are crucial in AI evaluation, including the target group, associated attribute, relationship characteristics, perceiving group, and context.
Scaling Cultural Resources for Improving Generative Models (2026.findings-eacl)

Copied to clipboard

Challenge: generative models have been known to have reduced performance in different global cultural contexts and languages.
Approach: They construct a pipeline to collect and contribute culturally salient, multilingual data . they argue such data can assess the state of the global applicability of generative AI models .
Outcome: The proposed pipeline can assess the state of the global applicability of our models and improve upon cross-cultural gaps.
Underspecification in Scene Description-to-Depiction Tasks (2022.aacl-main)

Copied to clipboard

Challenge: Recent text-to-image generation systems have demonstrated impressive capabilities . recent work focuses on generating images depicting scenes from scene descriptions .
Approach: They propose a conceptual framework to address implicitness, ambiguity and underspecification issues in multimodal image+text systems.
Outcome: The proposed framework addresses key challenges concerning textual and visual ambiguity and risks that may be amplified by ambiguous and underspecified elements.
Re-contextualizing Fairness in NLP: The Case of India (2022.aacl-main)

Copied to clipboard

Challenge: Recent research has revealed undesirable biases in NLP data and models . however, these efforts focus of social disparities in the West and are not directly portable to other geo-cultural contexts.
Approach: They propose a framework to re-contextualize NLP fairness research for the Indian context . they build resources for fairness evaluation in the Indian and delve deeper into social stereotypes for Region and Religion .
Outcome: The proposed framework can be generalized to other geo-cultural contexts.
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches for evaluating stereotypes have a noticeable lack of coverage of global identity groups and their associated stereotypes.
Approach: They propose to use a dataset to evaluate nationality-based stereotypes in T2I models across 135 nationalities to assess offensive stereotypes.
Outcome: The proposed dataset enables evaluation of known nationality-based stereotypes across 135 nationalities.
Towards Geo-Culturally Grounded LLM Generations (2025.acl-short)

Copied to clipboard

Challenge: Contemporary large language models (LLMs) are pretrained on huge corpora of natural language text and fine-tuned using human feedback to improve their quality.
Approach: They compare the performance of standard LLMs, LLM augmented with retrievals from a bespoke knowledge base and LLM with retrieval from . a web search on multiple cultural awareness benchmarks.
Outcome: The retrieval augmented generation and search grounding techniques improve LLMs' ability to display familiarity with various national cultures on cultural awareness benchmarks.
Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations (2022.tacl-1)

Copied to clipboard

Challenge: Annotators may systematically disagree with one another, reflecting their individual biases and values, especially in the case of subjective tasks such as detecting affect, aggression, and hate speech.
Approach: They propose to combine multi-annotator models with multi-task based approaches to resolve disagreements between annotations and derive single ground truth labels.
Outcome: The proposed model outperforms majority voting and averaging methods and estimates uncertainty in predictions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations