Rethinking Personality Assessment from Human-Agent Dialogues: Fewer Rounds May Be Better Than More (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing personality assessment datasets based on natural language do not consider interactivity. |
| Approach: | They propose to use a Chinese dataset to study the effects of different interaction rounds and agent personalities on personality assessment. |
| Outcome: | The proposed dataset contains 1260 interaction rounds between humans and agents with different personalities. |
Similar Papers
Beyond Self-Reports: Multi-Observer Agents for Personality Assessment in Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Self-report questionnaires are used to assess LLM personality traits, but they fail to capture behavioral nuances due to biases and meta-knowledge contamination. |
| Approach: | They propose a multi-observer framework for personality trait assessments in LLM agents that draws on informant-report methods in psychology. |
| Outcome: | The proposed framework combines multiple observers with a subject LLM agent to assess its Big Five personality traits. |
Can ChatGPT Assess Human Personalities? A General Evaluation Framework (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies study the virtual personalities of LLMs but rarely explore the possibility of analyzing human personalities via LLM. |
| Approach: | They propose to use Myers–Briggs Type Indicator (MBTI) tests to generate unbiased prompts and replace the subject in question statements to enable flexible queries and assessments. |
| Outcome: | The proposed framework enables LLMs to flexibly assess personalities of different groups of people. |
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture complexities of personalized task-oriented assistance. |
| Approach: | They propose a benchmark to evaluate personalization in task-oriented AI assistants . the benchmark features user profiles equipped with rich preferences and interaction histories . |
| Outcome: | The proposed benchmark features user profiles equipped with rich preferences and interaction histories . it also features a judge agent and user agent that employs the LLM-as-a-Judge paradigm . |
Be Selfish, But Wisely: Investigating the Impact of Agent Personality in Mixed-Motive Human-Agent Interactions (2023.emnlp-main)
Copied to clipboard
| Challenge: | A natural way to design a negotiation dialogue system is via self-play RL: train an agent that learns to maximize its performance by interacting with a simulated user that has been designed to imitate human-human dialogue data. |
| Approach: | They propose to use RL to train an agent that learns to maximize its performance by interacting with a simulated user that has been designed to imitate human-human dialogue data. |
| Outcome: | The proposed system fails to learn the value of compromise in a negotiation, which can lead to no agreements, and ultimately hurt the model's overall performance. |
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs can generate content that aligns with their assigned personality traits, but there is limited research on whether they consistently reflect specific personality traits. |
| Approach: | They propose to study the behavior of LLM-based agents which they refer to as LLM personas and simulate them to measure their personality traits. |
| Outcome: | The proposed model is based on the Big Five personality model and has been validated by human evaluations and automatic evaluations. |
Training Millions of Personalized Dialogue Agents (D18-1)
Copied to clipboard
| Challenge: | Current dialogue systems fail at being engaging for users when trained end-to-end without relying on proactive reengaging scripted strategies. |
| Approach: | They propose a dataset that provides 5 million personas and 700 million person-based dialogues. |
| Outcome: | The proposed dataset provides 5 million personas and 700 million person-based dialogues. |
Persona Dynamics: Unveiling the Impact of Persona Traits on Agents in Text-Based Games (2025.acl-long)
Copied to clipboard
| Challenge: | Text-based interactive environments have long presented formidable challenges for AI. |
| Approach: | They propose a method for projecting human personality traits onto agents to guide their behavior and integrate them into their policy-learning pipelines. |
| Outcome: | The proposed method induces personality in a text-based game agent by integrating personality profiles directly into the agent's policy-learning pipeline. |
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)
Copied to clipboard
| Challenge: | Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities. |
| Approach: | They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit. |
| Outcome: | The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries. |
Investigating the Personality Consistency in Quantized Role-Playing Dialogue Agents (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Using the Big Five personality traits model, we evaluate how stable assigned personalities are for Quantized Role-Playing Dialog Agents (QRPDA) during multi-turn interactions. |
| Approach: | They propose a non-parametric method to evaluate the stability of assigned personalities in quantized large language models (LLMs) for role-playing scenarios. |
| Outcome: | The proposed method shows that it maintains consistent personality traits in QRPDA, and it is more reliable in real-world applications. |
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation (2024.acl-long)
Copied to clipboard
| Challenge: | CharacterEval is a benchmark for comprehensive RPCA assessment in Chinese . authors show that Chinese LLMs exhibit more promising capabilities than GPT-4 in role-playing conversation. |
| Approach: | They propose a Chinese benchmark for comprehensive RPCA assessment . they use a dataset of Chinese role-playing dialogues and character profiles . |
| Outcome: | The proposed benchmark demonstrates that Chinese LLMs exhibit more promising capabilities than GPT-4 in Chinese role-playing conversation. |