Papers by Indira Sen

14 papers
Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection (2022.naacl-main)

Copied to clipboard

Challenge: sexism and hate speech detection models may be over-relying on core features . construct-driven CAD may induce models to ignore context in which core features are used .
Approach: They propose to use construct-driven and construct-agnostic CAD to reduce model bias . sexism and hate speech detection models are trained on counterfactually augmented data .
Outcome: Using a diverse set of CAD—construct-driven and construct-agnostic—reduces unintended bias.
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness (2025.findings-acl)

Copied to clipboard

Challenge: despite evidence of demographic bias, reports with whom they align best are hard to generalize or contradictory . confounders introduced in the annotation process account for more variation in alignment patterns than demographic traits .
Approach: They examine the alignment of large language models with human annotations in offensive language datasets.
Outcome: The results show that LLMs align better with human annotations than other models.
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)

Copied to clipboard

Challenge: Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests.
Approach: They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes .
Outcome: The proposed model improves model robustness on out-of-domain test sets and individual data points.
The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: persona prompting is increasingly used in large language models to simulate views of various sociodemographic groups.
Approach: They use open-source LLMs to study how persona prompts influence LLM simulations . they use role adoption formats and demographic priming strategies to study marginalized groups .
Outcome: The results show that the choice of demographic priming and role adoption strategy significantly impacts their portrayal.
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena.
Approach: They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality .
Outcome: The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias.
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Political biases in language models can affect performance in many applications . political biased models are often left-leaning, but are generally more left- leaning for instruction-tuned models .
Approach: They propose to use the Political Compass Test to measure political bias in language models . they use survey-based evaluation tools to test prompts and classify their political stances .
Outcome: The proposed model is based on the Political Compass Test, but is not scientifically valid.
Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) with chat interfaces are increasingly popular in various scientific fields, for a variety of tasks related to social science research questions.
Approach: They propose to use large language models to combine human and machine expertise to improve their models' performance.
Outcome: The proposed model performs better with co-created definitions than with expert-written definitions.
An Open Multilingual System for Scoring Readability of Wikipedia (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on the readability of Wikipedia have focused on English only and there are currently no systems supporting automatic readability assessment of the 300+ languages in Wikipedia.
Approach: They propose a multilingual model to assess Wikipedia's readability using a dataset spanning 14 languages.
Outcome: The proposed model outperforms existing models in a zero-shot scenario and is more accurate than previous benchmarks.
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories.
Approach: 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes .
Outcome: 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes .
Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality (2026.eacl-long)

Copied to clipboard

Challenge: Psychometric tests are increasingly used to assess psychological constructs in large language models (LLMs).
Approach: They evaluate the reliability and validity of human psychometric tests on 17 LLMs for three constructs: sexism, racism, and morality.
Outcome: The results show that the psychometric tests on 17 LLMs do not align, and in some cases negatively correlate with, model behavior in downstream tasks, indicating low ecological validity.
How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs? (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that models trained on CAD can learn cues in the dataset which are spuriously correlated with the construct.
Approach: They focus on sentiment, sexism, and hate speech as social constructs to investigate their effects on model performance.
Outcome: The proposed model generalizes better on out-of-domain datasets while relying less on spurious features.
On the Reliability and Validity of Detecting Approval of Political Actors in Tweets (2020.emnlp-main)

Copied to clipboard

Challenge: Social media sites have the potential to complement surveys that measure political opinions and, more specifically, political actors’ approval.
Approach: They propose to compare untargeted sentiment, targeted sentiment, and stance detection methods to a set of custom models trained on minimal custom data.
Outcome: The proposed methods have low generalizability on unseen and familiar targets, while low-resource custom models are more robust.
Language Identification and Named Entity Recognition in Hinglish Code Mixed Tweets (P18-3)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is an important text analysis task . code-mixing occurs when lexical items and grammatical features from two languages appear in one sentence .
Approach: They propose to use language identifiers, parts-of-speech tags and chunkers to analyze code-mixed data.
Outcome: The proposed method outperforms the best baseline by 33.18%.
Neural network embeddings recover value dimensions from psychometric survey items on par with human data (2026.findings-eacl)

Copied to clipboard

Challenge: Embedings from large language models can recover structure of human values . quantitative analysis reveals that SQuID addresses the challenge of obtaining negative correlations between dimensions without domain-specific fine-tuning or training data reannotation.
Approach: They propose to use questionnaire item embeddings to recover human values from PVQ-RR . their results have implications for psychometrics and social science research .
Outcome: The proposed method explains 55% variance in dimension-dimension similarities compared to human data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations