Papers by Indira Sen
Counterfactually Augmented Data and Unintended Bias: The Case of Sexism and Hate Speech Detection (2022.naacl-main)
Copied to clipboard
| Challenge: | sexism and hate speech detection models may be over-relying on core features . construct-driven CAD may induce models to ignore context in which core features are used . |
| Approach: | They propose to use construct-driven and construct-agnostic CAD to reduce model bias . sexism and hate speech detection models are trained on counterfactually augmented data . |
| Outcome: | Using a diverse set of CAD—construct-driven and construct-agnostic—reduces unintended bias. |
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness (2025.findings-acl)
Copied to clipboard
| Challenge: | despite evidence of demographic bias, reports with whom they align best are hard to generalize or contradictory . confounders introduced in the annotation process account for more variation in alignment patterns than demographic traits . |
| Approach: | They examine the alignment of large language models with human annotations in offensive language datasets. |
| Outcome: | The results show that LLMs align better with human annotations than other models. |
People Make Better Edits: Measuring the Efficacy of LLM-Generated Counterfactually Augmented Data for Harmful Language Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | Past work has shown that counterfactually augmented data (CADs) can improve models' performance on out-of-domain tests. |
| Approach: | They use Polyjuice, ChatGPT, and Flan-T5 to automatically generate CADs . they find that CAD generates a model that flips the original label with minimal changes . |
| Outcome: | The proposed model improves model robustness on out-of-domain test sets and individual data points. |
The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | persona prompting is increasingly used in large language models to simulate views of various sociodemographic groups. |
| Approach: | They use open-source LLMs to study how persona prompts influence LLM simulations . they use role adoption formats and demographic priming strategies to study marginalized groups . |
| Outcome: | The results show that the choice of demographic priming and role adoption strategy significantly impacts their portrayal. |
Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation (2026.eacl-long)
Copied to clipboard
Bolei Ma, Yong Cao, Indira Sen, Anna-Carolina Haensch, Frauke Kreuter, Barbara Plank, Daniel Hershcovich
| Challenge: | Large Language Models (LLMs) are increasingly used to simulate public opinion and other social phenomena. |
| Approach: | They argue that open-endedness is essential for realistic social simulations . they argue that it captures expressiveness and individuality . |
| Outcome: | The proposed frameworks can improve measurement and design, support exploration of unanticipated views, and reduce researcher-imposed directive bias. |
Only a Little to the Left: A Theory-grounded Measure of Political Bias in Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Political biases in language models can affect performance in many applications . political biased models are often left-leaning, but are generally more left- leaning for instruction-tuned models . |
| Approach: | They propose to use the Political Compass Test to measure political bias in language models . they use survey-based evaluation tools to test prompts and classify their political stances . |
| Outcome: | The proposed model is based on the Political Compass Test, but is not scientifically valid. |
Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) with chat interfaces are increasingly popular in various scientific fields, for a variety of tasks related to social science research questions. |
| Approach: | They propose to use large language models to combine human and machine expertise to improve their models' performance. |
| Outcome: | The proposed model performs better with co-created definitions than with expert-written definitions. |
An Open Multilingual System for Scoring Readability of Wikipedia (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies on the readability of Wikipedia have focused on English only and there are currently no systems supporting automatic readability assessment of the 300+ languages in Wikipedia. |
| Approach: | They propose a multilingual model to assess Wikipedia's readability using a dataset spanning 14 languages. |
| Outcome: | The proposed model outperforms existing models in a zero-shot scenario and is more accurate than previous benchmarks. |
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories. |
| Approach: | 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes . |
| Outcome: | 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes . |
Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality (2026.eacl-long)
Copied to clipboard
| Challenge: | Psychometric tests are increasingly used to assess psychological constructs in large language models (LLMs). |
| Approach: | They evaluate the reliability and validity of human psychometric tests on 17 LLMs for three constructs: sexism, racism, and morality. |
| Outcome: | The results show that the psychometric tests on 17 LLMs do not align, and in some cases negatively correlate with, model behavior in downstream tasks, indicating low ecological validity. |
How Does Counterfactually Augmented Data Impact Models for Social Computing Constructs? (2021.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have shown that models trained on CAD can learn cues in the dataset which are spuriously correlated with the construct. |
| Approach: | They focus on sentiment, sexism, and hate speech as social constructs to investigate their effects on model performance. |
| Outcome: | The proposed model generalizes better on out-of-domain datasets while relying less on spurious features. |
On the Reliability and Validity of Detecting Approval of Political Actors in Tweets (2020.emnlp-main)
Copied to clipboard
| Challenge: | Social media sites have the potential to complement surveys that measure political opinions and, more specifically, political actors’ approval. |
| Approach: | They propose to compare untargeted sentiment, targeted sentiment, and stance detection methods to a set of custom models trained on minimal custom data. |
| Outcome: | The proposed methods have low generalizability on unseen and familiar targets, while low-resource custom models are more robust. |
Language Identification and Named Entity Recognition in Hinglish Code Mixed Tweets (P18-3)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is an important text analysis task . code-mixing occurs when lexical items and grammatical features from two languages appear in one sentence . |
| Approach: | They propose to use language identifiers, parts-of-speech tags and chunkers to analyze code-mixed data. |
| Outcome: | The proposed method outperforms the best baseline by 33.18%. |
Neural network embeddings recover value dimensions from psychometric survey items on par with human data (2026.findings-eacl)
Copied to clipboard
| Challenge: | Embedings from large language models can recover structure of human values . quantitative analysis reveals that SQuID addresses the challenge of obtaining negative correlations between dimensions without domain-specific fine-tuning or training data reannotation. |
| Approach: | They propose to use questionnaire item embeddings to recover human values from PVQ-RR . their results have implications for psychometrics and social science research . |
| Outcome: | The proposed method explains 55% variance in dimension-dimension similarities compared to human data. |