Challenge: Persona-driven large language models (LLMs) are increasingly used in computational social science, yet their validity critically depends on the fidelity of the underlying personas.
Approach: They propose a persona-generation framework that grounds LLM-generated personas in real social-media posts while delegating narrative construction to language models.
Outcome: The proposed framework outperforms state-of-the-art methods for most demographics across different dimensions while maintaining interaction graph structure among personas grounded in real social network users.

Similar Papers

GRAVITY: A Framework for Personalized Text Generation via Profile-Grounded Synthetic Preferences (2026.eacl-long)

Copied to clipboard

Challenge: Personalization in LLMs often relies on costly human feedback or interaction logs, limiting scalability and neglecting deeper user attributes.
Approach: They propose a framework for generating synthetic, profile-grounded preference data that captures users’ interests, values, beliefs, and personality traits.
Outcome: The proposed framework improves on book descriptions for 400 Amazon users across multiple cultures, with user studies showing that outputs are preferred over 86% of the time.
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing static benchmarks for harmful content detection face limitations in scalability and diversity.
Approach: They propose a framework for synthesizing harmful content using persona-guided large language model agents.
Outcome: The proposed framework achieves a high success rate in harmful generation tests across multiple detection systems.
PERSONA: A Reproducible Testbed for Pluralistic Alignment (2025.coling-main)

Copied to clipboard

Challenge: Currently, preference optimization approaches fail to capture the plurality of user opinions . Currently used methods do not account for the pluralities of users and difference of opinion .
Approach: They propose a reproducible test bed to evaluate pluralistic alignment of language models . they generate user profiles from census data and use a large-scale evaluation dataset .
Outcome: The proposed model improves pluralistic alignment of language models with diverse user values . it generates a large-scale evaluation dataset with 317,200 feedback pairs .
Faithful Persona-based Conversational Dataset Generation with Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing datasets for training conversational AI models do not sufficiently model their users.
Approach: They propose a generator-critic architecture framework to expand the initial dataset while improving the quality of its conversations.
Outcome: The proposed framework expands the initial dataset while improving the quality of its conversations.
A Structured Clustering Approach for Inducing Media Narratives (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to modeling media narratives miss subtle narrative patterns through coarse-grained analysis or require domain-specific taxonomies that limit scalability.
Approach: They propose a framework for inducing rich narrative schemas by jointly modeling events and characters via structured clustering.
Outcome: The proposed framework produces explainable narrative schemas that align with established framing theory while scaling to large corpora without exhaustive manual annotation.
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities.
Approach: They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit.
Outcome: The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries.
HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents (2026.findings-acl)

Copied to clipboard

Challenge: ad-hoc prompting and hand-crafted profiles with limited control over educational theory and population distributions are often used for student personas.
Approach: They propose a framework that generates theory-aligned, quota-controlled personas . they factorize each persona into a theory-anchored educational schema .
Outcome: HACHIMI generates theory-aligned, quota-controlled personas for grades 1-12 . results show near-perfect schema validity, accurate quots, and substantial diversity .
Grounding in social media: An approach to building a chit-chat dialogue model (2022.naacl-srw)

Copied to clipboard

Challenge: Existing open-domain dialogue models fail to capture and utilize external knowledge, leading to repetitive or generic responses to unseen utterances.
Approach: They propose to use social media comments to improve the raw conversation ability of open-domain dialogue systems.
Outcome: The proposed model improves the raw conversation ability of open-domain dialogue systems by mimicking human responses through casual interactions found on social media.
Evaluating Large Language Model Biases in Persona-Steered Generation (2024.findings-acl)

Copied to clipboard

Challenge: a recent wave of powerful new large language models has raised concerns that their expressed opinions may be biased towards certain political, national or moral viewpoints.
Approach: They define an incongruous persona as a persona with multiple traits where one trait makes its other traits less likely in human survey data.
Outcome: The results show that LLMs are less steerable towards incongruous personas than congruous ones . the models that are fine-tuned with RLHF are more steerable, especially towards stances associated with political liberals and women .
SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Recent calls for pluralistic alignment of Large Language Models encourage adapting models to diverse user preferences.
Approach: They propose a method to induce synthetic user personas from user interactions for personalized reward modeling.
Outcome: The proposed approach improves LLM-as-a-judge accuracy by 4.4% on Chatbot Arena.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations