Challenge: Existing personality assessment datasets based on natural language do not consider interactivity.
Approach: They propose to use a Chinese dataset to study the effects of different interaction rounds and agent personalities on personality assessment.
Outcome: The proposed dataset contains 1260 interaction rounds between humans and agents with different personalities.

Similar Papers

Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Self-report questionnaires are used to assess LLM personality traits, but they fail to capture behavioral nuances due to biases and meta-knowledge contamination.
Approach: They propose a multi-observer framework for personality trait assessments in LLM agents that draws on informant-report methods in psychology.
Outcome: The proposed framework combines multiple observers with a subject LLM agent to assess its Big Five personality traits.
Can ChatGPT Assess Human Personalities? A General Evaluation Framework (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies study the virtual personalities of LLMs but rarely explore the possibility of analyzing human personalities via LLM.
Approach: They propose to use Myers–Briggs Type Indicator (MBTI) tests to generate unbiased prompts and replace the subject in question statements to enable flexible queries and assessments.
Outcome: The proposed framework enables LLMs to flexibly assess personalities of different groups of people.
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants (2025.findings-acl)

Copied to clipboard

Challenge: Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture complexities of personalized task-oriented assistance.
Approach: They propose a benchmark to evaluate personalization in task-oriented AI assistants . the benchmark features user profiles equipped with rich preferences and interaction histories .
Outcome: The proposed benchmark features user profiles equipped with rich preferences and interaction histories . it also features a judge agent and user agent that employs the LLM-as-a-Judge paradigm .
Be Selfish, But Wisely: Investigating the Impact of Agent Personality in Mixed-Motive Human-Agent Interactions (2023.emnlp-main)

Copied to clipboard

Challenge: A natural way to design a negotiation dialogue system is via self-play RL: train an agent that learns to maximize its performance by interacting with a simulated user that has been designed to imitate human-human dialogue data.
Approach: They propose to use RL to train an agent that learns to maximize its performance by interacting with a simulated user that has been designed to imitate human-human dialogue data.
Outcome: The proposed system fails to learn the value of compromise in a negotiation, which can lead to no agreements, and ultimately hurt the model's overall performance.
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies have shown that LLMs can generate content that aligns with their assigned personality traits, but there is limited research on whether they consistently reflect specific personality traits.
Approach: They propose to study the behavior of LLM-based agents which they refer to as LLM personas and simulate them to measure their personality traits.
Outcome: The proposed model is based on the Big Five personality model and has been validated by human evaluations and automatic evaluations.
Training Millions of Personalized Dialogue Agents (D18-1)

Copied to clipboard

Challenge: Current dialogue systems fail at being engaging for users when trained end-to-end without relying on proactive reengaging scripted strategies.
Approach: They propose a dataset that provides 5 million personas and 700 million person-based dialogues.
Outcome: The proposed dataset provides 5 million personas and 700 million person-based dialogues.
Persona Dynamics: Unveiling the Impact of Persona Traits on Agents in Text-Based Games (2025.acl-long)

Copied to clipboard

Challenge: Text-based interactive environments have long presented formidable challenges for AI.
Approach: They propose a method for projecting human personality traits onto agents to guide their behavior and integrate them into their policy-learning pipelines.
Outcome: The proposed method induces personality in a text-based game agent by integrating personality profiles directly into the agent's policy-learning pipeline.
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities.
Approach: They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit.
Outcome: The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries.
Investigating the Personality Consistency in Quantized Role-Playing Dialogue Agents (2024.emnlp-industry)

Copied to clipboard

Challenge: Using the Big Five personality traits model, we evaluate how stable assigned personalities are for Quantized Role-Playing Dialog Agents (QRPDA) during multi-turn interactions.
Approach: They propose a non-parametric method to evaluate the stability of assigned personalities in quantized large language models (LLMs) for role-playing scenarios.
Outcome: The proposed method shows that it maintains consistent personality traits in QRPDA, and it is more reliable in real-world applications.
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation (2024.acl-long)

Copied to clipboard

Challenge: CharacterEval is a benchmark for comprehensive RPCA assessment in Chinese . authors show that Chinese LLMs exhibit more promising capabilities than GPT-4 in role-playing conversation.
Approach: They propose a Chinese benchmark for comprehensive RPCA assessment . they use a dataset of Chinese role-playing dialogues and character profiles .
Outcome: The proposed benchmark demonstrates that Chinese LLMs exhibit more promising capabilities than GPT-4 in Chinese role-playing conversation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations