Synthia: Scalable Grounded Persona Generation from Social Media Data (2026.acl-long)
Copied to clipboard
| Challenge: | Persona-driven large language models (LLMs) are increasingly used in computational social science, yet their validity critically depends on the fidelity of the underlying personas. |
| Approach: | They propose a persona-generation framework that grounds LLM-generated personas in real social-media posts while delegating narrative construction to language models. |
| Outcome: | The proposed framework outperforms state-of-the-art methods for most demographics across different dimensions while maintaining interaction graph structure among personas grounded in real social network users. |
Similar Papers
GRAVITY: A Framework for Personalized Text Generation via Profile-Grounded Synthetic Preferences (2026.eacl-long)
Copied to clipboard
| Challenge: | Personalization in LLMs often relies on costly human feedback or interaction logs, limiting scalability and neglecting deeper user attributes. |
| Approach: | They propose a framework for generating synthetic, profile-grounded preference data that captures users’ interests, values, beliefs, and personality traits. |
| Outcome: | The proposed framework improves on book descriptions for 400 Amazon users across multiple cultures, with user studies showing that outputs are preferred over 86% of the time. |
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | Existing static benchmarks for harmful content detection face limitations in scalability and diversity. |
| Approach: | They propose a framework for synthesizing harmful content using persona-guided large language model agents. |
| Outcome: | The proposed framework achieves a high success rate in harmful generation tests across multiple detection systems. |
PERSONA: A Reproducible Testbed for Pluralistic Alignment (2025.coling-main)
Copied to clipboard
| Challenge: | Currently, preference optimization approaches fail to capture the plurality of user opinions . Currently used methods do not account for the pluralities of users and difference of opinion . |
| Approach: | They propose a reproducible test bed to evaluate pluralistic alignment of language models . they generate user profiles from census data and use a large-scale evaluation dataset . |
| Outcome: | The proposed model improves pluralistic alignment of language models with diverse user values . it generates a large-scale evaluation dataset with 317,200 feedback pairs . |
Faithful Persona-based Conversational Dataset Generation with Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for training conversational AI models do not sufficiently model their users. |
| Approach: | They propose a generator-critic architecture framework to expand the initial dataset while improving the quality of its conversations. |
| Outcome: | The proposed framework expands the initial dataset while improving the quality of its conversations. |
A Structured Clustering Approach for Inducing Media Narratives (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to modeling media narratives miss subtle narrative patterns through coarse-grained analysis or require domain-specific taxonomies that limit scalability. |
| Approach: | They propose a framework for inducing rich narrative schemas by jointly modeling events and characters via structured clustering. |
| Outcome: | The proposed framework produces explainable narrative schemas that align with established framing theory while scaling to large corpora without exhaustive manual annotation. |
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)
Copied to clipboard
| Challenge: | Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities. |
| Approach: | They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit. |
| Outcome: | The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries. |
HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | ad-hoc prompting and hand-crafted profiles with limited control over educational theory and population distributions are often used for student personas. |
| Approach: | They propose a framework that generates theory-aligned, quota-controlled personas . they factorize each persona into a theory-anchored educational schema . |
| Outcome: | HACHIMI generates theory-aligned, quota-controlled personas for grades 1-12 . results show near-perfect schema validity, accurate quots, and substantial diversity . |
Grounding in social media: An approach to building a chit-chat dialogue model (2022.naacl-srw)
Copied to clipboard
| Challenge: | Existing open-domain dialogue models fail to capture and utilize external knowledge, leading to repetitive or generic responses to unseen utterances. |
| Approach: | They propose to use social media comments to improve the raw conversation ability of open-domain dialogue systems. |
| Outcome: | The proposed model improves the raw conversation ability of open-domain dialogue systems by mimicking human responses through casual interactions found on social media. |
Evaluating Large Language Model Biases in Persona-Steered Generation (2024.findings-acl)
Copied to clipboard
| Challenge: | a recent wave of powerful new large language models has raised concerns that their expressed opinions may be biased towards certain political, national or moral viewpoints. |
| Approach: | They define an incongruous persona as a persona with multiple traits where one trait makes its other traits less likely in human survey data. |
| Outcome: | The results show that LLMs are less steerable towards incongruous personas than congruous ones . the models that are fine-tuned with RLHF are more steerable, especially towards stances associated with political liberals and women . |
SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs (2025.acl-long)
Copied to clipboard
| Challenge: | Recent calls for pluralistic alignment of Large Language Models encourage adapting models to diverse user preferences. |
| Approach: | They propose a method to induce synthetic user personas from user interactions for personalized reward modeling. |
| Outcome: | The proposed approach improves LLM-as-a-judge accuracy by 4.4% on Chatbot Arena. |