Challenge: Recent advances in Large Language Models (LLMs) have led to their adaptation as conversational agents.
Approach: They propose a new benchmark that uses 8K multi-choice questions to assess the personality of Large Language Models.
Outcome: The proposed personality test outperforms existing personality tests for LLMs in reliability and validity.

Similar Papers

PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits (2024.findings-naacl)

Copied to clipboard

Challenge: Recent studies have shown that LLMs can generate content that aligns with their assigned personality traits, but there is limited research on whether they consistently reflect specific personality traits.
Approach: They propose to study the behavior of LLM-based agents which they refer to as LLM personas and simulate them to measure their personality traits.
Outcome: The proposed model is based on the Big Five personality model and has been validated by human evaluations and automatic evaluations.
On the Reliability of Psychological Scales on Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has focused on examining Large Language Models’ characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics.
Approach: They propose to examine the reliability of personality tests to LLMs by using psychological scales.
Outcome: The proposed model can represent diverse personalities with specific prompt instructions.
You don’t need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are popular for research in social sciences . currently, prompting LLMs is insufficient to accurately and reliably capture model perceptions, and we discuss potential alternatives to improve this.
Approach: They construct a dataset that contains 693 questions encompassing 39 different instruments of persona measurement on 115 persona axes and a set of questions containing minor variations.
Outcome: The proposed model can generate answers and negate statements in a consistent and robust manner.
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have explored personality evaluation of LLMs, but they largely overlook the interplay between culture and personality.
Approach: They propose a large-scale benchmark for evaluating LLMs’ personality expression in culturally grounded, behaviorally rich contexts.
Outcome: The proposed benchmark improves alignment with country-specific human personality distributions and elicits more expressive, culturally coherent outputs compared to existing benchmarks.
Modeling, Evaluating, and Embodying Personality in LLMs: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey provides a comprehensive overview of the LLM-driven personality scenario.
Approach: This survey provides a comprehensive overview of the LLM-driven personality scenario.
Outcome: The proposed taxonomy analyzes the limitations of existing methods and identifies key research gaps.
Can ChatGPT Assess Human Personalities? A General Evaluation Framework (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies study the virtual personalities of LLMs but rarely explore the possibility of analyzing human personalities via LLM.
Approach: They propose to use Myers–Briggs Type Indicator (MBTI) tests to generate unbiased prompts and replace the subject in question statements to enable flexible queries and assessments.
Outcome: The proposed framework enables LLMs to flexibly assess personalities of different groups of people.
Can LLM Agents Maintain a Persona in Discourse? (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models are often subjected to context-shifting behaviour, resulting in a lack of consistent and interpretable personality-aligned interactions.
Approach: They propose to use two conversation agents to generate a discourse with an assigned personality from the OCEAN framework and then use multiple judge agents to infer original traits.
Outcome: The proposed model is based on two conversation agents with a personality assigned from the OCEAN framework and then multiple judge agents to infer the original traits assigned.
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Personalized Large Language Models are increasingly used in diverse applications . prior research examined how well LLMs adhere to predefined personas in writing style . inconsistent responses are influenced by multiple factors, including the assigned persona, stereotypes, and model design choices.
Approach: They propose a standardized framework to analyze consistency in persona-assigned LLMs.
Outcome: The proposed framework evaluates personas across multiple tasks and runs.
Decoding LLM Personality Measurement: Forced-Choice vs. Likert (2025.findings-acl)

Copied to clipboard

Challenge: Recent research has focused on investigating the psychological characteristics of Large Language Models (LLMs), emphasizing the importance of comprehending their behavioral traits.
Approach: They evaluated six Large Language Models: Llama-3.1-8B, GLM-4-9B, Claude-3.5-sonnet, and Deepseek-V3 and used the forced-choice test to assess their personality traits.
Outcome: The forced-choice test is more reliable and more accurate than the likert scale and forced-CHOICE test results for LLMs' Big Five personality scores.
Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly integrated into everyday workflows . a recent study found that LLMs exhibit distinct personality-like traits that affect user engagement .
Approach: They evaluated 32 LLM users for four collaborative tasks and found significant preferences . they found that rationalists preferred GPT-4, while idealists favored Claude 3.5 .
Outcome: The results show that users with different personality traits prefer certain LLMs over others.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations