CAPE: Context-Aware Personality Evaluation Framework for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies use a context-free approach to assess humans . existing studies use the Disney World test, which ignores real-world applications . |
| Approach: | They propose a framework to assess personality traits in large language models . they use conversational history to quantify the consistency of LLM responses . |
| Outcome: | The proposed framework improves consistency of responses in large language models . it also shows that conversational history enhances consistency and personality shifts . |
Similar Papers
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits (2024.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have shown that LLMs can generate content that aligns with their assigned personality traits, but there is limited research on whether they consistently reflect specific personality traits. |
| Approach: | They propose to study the behavior of LLM-based agents which they refer to as LLM personas and simulate them to measure their personality traits. |
| Outcome: | The proposed model is based on the Big Five personality model and has been validated by human evaluations and automatic evaluations. |
Can ChatGPT Assess Human Personalities? A General Evaluation Framework (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies study the virtual personalities of LLMs but rarely explore the possibility of analyzing human personalities via LLM. |
| Approach: | They propose to use Myers–Briggs Type Indicator (MBTI) tests to generate unbiased prompts and replace the subject in question statements to enable flexible queries and assessments. |
| Outcome: | The proposed framework enables LLMs to flexibly assess personalities of different groups of people. |
On the Reliability of Psychological Scales on Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent research has focused on examining Large Language Models’ characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics. |
| Approach: | They propose to examine the reliability of personality tests to LLMs by using psychological scales. |
| Outcome: | The proposed model can represent diverse personalities with specific prompt instructions. |
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics (2025.findings-naacl)
Copied to clipboard
Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, Jinyoung Yeo, Youngjae Yu
| Challenge: | Recent advances in Large Language Models (LLMs) have led to their adaptation as conversational agents. |
| Approach: | They propose a new benchmark that uses 8K multi-choice questions to assess the personality of Large Language Models. |
| Outcome: | The proposed personality test outperforms existing personality tests for LLMs in reliability and validity. |
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have explored personality evaluation of LLMs, but they largely overlook the interplay between culture and personality. |
| Approach: | They propose a large-scale benchmark for evaluating LLMs’ personality expression in culturally grounded, behaviorally rich contexts. |
| Outcome: | The proposed benchmark improves alignment with country-specific human personality distributions and elicits more expressive, culturally coherent outputs compared to existing benchmarks. |
ChildEval:WHEN LARGE LANGUAGE MODELS MEET CHILDREN’S PERSONALITIES (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved remarkable success in effectively understanding and generating human language, leading to a revolutionary era in LLMs. |
| Approach: | They propose a benchmark to evaluate LLMs' ability to infer and follow child-centered preferences in long-context conversations. |
| Outcome: | The proposed benchmark spans five top-level and fourteen sub-level categories covering children’s daily lives and development. |
Modeling, Evaluating, and Embodying Personality in LLMs: A Survey (2025.findings-emnlp)
Copied to clipboard
Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Pedro Schindler Freire Brasil Ribeiro, Rafael Teixeira Sousa, Arlindo Rodrigues Galvão Filho
| Challenge: | This survey provides a comprehensive overview of the LLM-driven personality scenario. |
| Approach: | This survey provides a comprehensive overview of the LLM-driven personality scenario. |
| Outcome: | The proposed taxonomy analyzes the limitations of existing methods and identifies key research gaps. |
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for embedding human personality traits into LLMs are limited by realism and validity issues. |
| Approach: | They propose to use a large-scale dataset to embed human personality traits into LLMs . they use supervised fine-tuning and direct preference optimization to train LLM models . |
| Outcome: | The proposed methods outperform prompting on personality assessments and IPIP-NEO, and show higher conscientiousness, agreeableness, lower extraversion, and lower neuroticism on reasoning tasks. |
Manipulating the Perceived Personality Traits of Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Psychology research has long explored aspects of human personality like extroversion, agreeableness and emotional stability, three of the personality traits that make up the ‘Big Five’. |
| Approach: | They propose to use text generated from large language models to evaluate perceived personality traits and to frame them as tools for controlling personas in dialog systems. |
| Outcome: | The proposed models predict personality traits in different contexts and can be manipulated in a predictable way. |
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results. |
| Approach: | They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
| Outcome: | The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |