The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | persona prompting is increasingly used in large language models to simulate views of various sociodemographic groups. |
| Approach: | They use open-source LLMs to study how persona prompts influence LLM simulations . they use role adoption formats and demographic priming strategies to study marginalized groups . |
| Outcome: | The results show that the choice of demographic priming and role adoption strategy significantly impacts their portrayal. |
Similar Papers
Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Persona-prompting is a growing strategy to personalize outputs, but its impact on how LLMs represent social groups remains underexplored. |
| Approach: | They investigate whether persona-prompting leads to different levels of linguistic abstraction . they compare 11 persona driven responses to those of a generic AI assistant . |
| Outcome: | The proposed method can be used to personalize outputs, but its impact on how LLMs represent social groups remains underexplored. |
Quantifying the Persona Effect in LLM Simulations (2024.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown remarkable promise in simulating human language and behavior. |
| Approach: | They investigate how integrating persona variables—demographic, social, and behavioral factors—impacts LLMs’ ability to simulate diverse perspectives. |
| Outcome: | The proposed model improves on a zero-shot model with persona prompting. |
One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization (2026.acl-long)
Copied to clipboard
| Challenge: | Prior work has used personas to study biases by relying on a single cue to prompt a persona, such as user names or explicit attribute mentions. |
| Approach: | They compare six commonly used personacues across seven open and proprietary LLMs on four writing and advice tasks. |
| Outcome: | The proposed model is based on a persona, a synthetic user profile defined by specific attributes, defined by gender or race. |
Persona Prompting as a Lens on LLM Social Reasoning (2026.eacl-long)
Copied to clipboard
Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann, Vera Schmitt, Nils Feldhus
| Challenge: | Persona prompting (PP) is increasingly used to steer large language models towards user-specific generation, but its effect on rationales remains underexplored. |
| Approach: | They examine how LLM-generated rationales vary when conditioned on different demographic personas . they use word-level rationale annotations to measure agreement with human annotations based on PP . |
| Outcome: | The proposed model improves classification on the most subjective task, but fails to align with real-world demographic counterparts. |
Sociodemographic Prompting is Not Yet an Effective Approach for Simulating Subjective Judgments with LLMs (2025.naacl-short)
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely used to simulate human responses, but their ability to account for demographic differences in subjective tasks remains uncertain. |
| Approach: | They evaluate large language models' ability to understand demographic differences in two subjective judgment tasks: politeness and offensiveness. |
| Outcome: | The proposed models perform better in politeness and offensiveness tasks, while sociodemographic prompting does not improve and worsens their ability to perceive language from sub-populations. |
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods to measure stereotypes in large language models rely on manual templates or natural sentences that contain stereotypes. |
| Approach: | They propose a prompt-based method to measure stereotypes in large language models . they use natural language descriptions of the target demographic group alongside unmarked defaults . |
| Outcome: | The proposed method detects that portrayals contain higher rates of racial stereotypes than human-written portrayals. |
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing studies on sociodemographic prompting have not explored the effectiveness of this technique. |
| Approach: | They propose to use sociodemographic prompting to steer models towards answers that humans with specific sociodemography would give. |
| Outcome: | The proposed technique can improve zero-shot learning by focusing on human sociodemographic profiles. |
Reading Between the Prompts: How Stereotypes Shape LLM’s Implicit Personalization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups. |
| Approach: | They analyze LLMs' latent user representations through both model internals and generated answers to targeted user questions. |
| Outcome: | The proposed models infer demographic attributes based on stereotypical signals, which persists even when the user explicitly identifies with a different demographic group. |
You don’t need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments (2024.naacl-long)
Copied to clipboard
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, David Jurgens
| Challenge: | Large Language Models (LLMs) are popular for research in social sciences . currently, prompting LLMs is insufficient to accurately and reliably capture model perceptions, and we discuss potential alternatives to improve this. |
| Approach: | They construct a dataset that contains 693 questions encompassing 39 different instruments of persona measurement on 115 persona axes and a set of questions containing minor variations. |
| Outcome: | The proposed model can generate answers and negate statements in a consistent and robust manner. |
When ”A Helpful Assistant” Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Commercial AI systems often define the role of the LLM in system prompts. |
| Approach: | They conduct a systematic evaluation of personas in system prompts by adding 162 roles covering 6 types of interpersonal relationships and 8 domains of expertise. |
| Outcome: | The proposed model does not improve performance in the system prompt setting where no persona is added. |