Fingerprinting LLMs through Survey Item Factor Correlation: A Case Study on Humor Style Questionnaire (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for evaluating LLMs focus on output accuracy, faithfulness, or alignment with human preferences, but these metrics do not capture fundamental differences in how models internally represent and relate psychological constructs. |
| Approach: | They propose to “fingerprint” LLMs through factor correlation patterns on standardized psychological assessments to deepen understanding of LLM's constructs representation. |
| Outcome: | The proposed method shows that LLMs represent constructs differently than humans . it also shows that no LLM recovers the constructs of the Humor Style Questionnaire . |
Similar Papers
Uncovering Factor-Level Preference to Improve Human-Model Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs. |
| Approach: | They propose a framework to uncover and measure factor-level preference alignment of humans and large language models (LLMs) |
| Outcome: | The proposed framework uncovers and measures factor-level preference alignment of humans and large language models. |
Modeling, Evaluating, and Embodying Personality in LLMs: A Survey (2025.findings-emnlp)
Copied to clipboard
Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Pedro Schindler Freire Brasil Ribeiro, Rafael Teixeira Sousa, Arlindo Rodrigues Galvão Filho
| Challenge: | This survey provides a comprehensive overview of the LLM-driven personality scenario. |
| Approach: | This survey provides a comprehensive overview of the LLM-driven personality scenario. |
| Outcome: | The proposed taxonomy analyzes the limitations of existing methods and identifies key research gaps. |
Quantifying Data Contamination in Psychometric Evaluations of LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing studies have raised concerns about data contamination from psychometric inventories . however, there is no systematic attempt to quantify the extent of data contamination . |
| Approach: | They propose a framework to measure data contamination in psychometric evaluations of Large Language Models by item memorization, evaluation memorisation and target score matching. |
| Outcome: | The proposed framework evaluates item memorization, evaluation memorisation, and target score matching in 21 models from major families and four widely used psychometric inventories. |
Assessment and manipulation of latent constructs in pre-trained language models using psychometric scales (2025.acl-long)
Copied to clipboard
Maor Reuben, Ortal Slobodin, Idan-Chaim Cohen, Aviad Elyashar, Orna Braun-Lewensohn, Odeya Cohen, Rami Puzis
| Challenge: | a recent study suggests that language models may be tricked into answering psychometric questionnaires, but they cannot be assessed because of inadequate psychometric methods. |
| Approach: | They propose to re-form standard psychological questionnaires into natural language inference prompts and a code library to support the psychometric assessment of arbitrary models. |
| Outcome: | The proposed model can be reformulated into natural language inference prompts and a code library to support the psychometric assessment of arbitrary models. |
Pun Unintended: LLMs and the Illusion of Humor Understanding (2025.emnlp-main)
Copied to clipboard
Alessandro Zangari, Matteo Marcuzzo, Andrea Albarelli, Mohammad Taher Pilehvar, Jose Camacho-Collados
| Challenge: | Existing models for pun detection lack nuanced grasp typical of human interpretation. |
| Approach: | They analyze existing pun detection benchmarks and human evaluation across recent LLMs to find subtle changes in puns that mislead LLM. |
| Outcome: | The proposed models lack the nuance typical of human interpretation and lack the depth of their analysis to detect puns. |
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics (2025.findings-naacl)
Copied to clipboard
Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, Jinyoung Yeo, Youngjae Yu
| Challenge: | Recent advances in Large Language Models (LLMs) have led to their adaptation as conversational agents. |
| Approach: | They propose a new benchmark that uses 8K multi-choice questions to assess the personality of Large Language Models. |
| Outcome: | The proposed personality test outperforms existing personality tests for LLMs in reliability and validity. |
Can LLMs Express Personality Across Cultures? Introducing CulturalPersonas for Evaluating Trait Alignment (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have explored personality evaluation of LLMs, but they largely overlook the interplay between culture and personality. |
| Approach: | They propose a large-scale benchmark for evaluating LLMs’ personality expression in culturally grounded, behaviorally rich contexts. |
| Outcome: | The proposed benchmark improves alignment with country-specific human personality distributions and elicits more expressive, culturally coherent outputs compared to existing benchmarks. |
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability. |
| Approach: | They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria. |
| Outcome: | The proposed system is based on 11 common aspects with different evaluation criteria. |
An Empirical Analysis of the Writing Styles of Persona-Assigned LLMs (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent efforts to "personalize" large language models by assigning them specific personas are limited by current knowledge of how well they perform. |
| Approach: | They use a style embedding model to analyze writing styles of persona-assigned LLMs . they find significant style differences between personas using Kullback-Leibler divergence . |
| Outcome: | The proposed model shows significant differences in writing styles among personas across socio-demographic groups. |
Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models can lead to undesired consequences when misaligned with human values . previous studies have shown misalignment of LLMs with human value using expert-designed or agent-based emulated bias scenarios . |
| Approach: | They investigate whether large language models (LLMs) are misaligned with human values . they find no significant differences in understanding of HVSB between LLMs . |
| Outcome: | The results show that large language models do not have lower misalignment rates and attack success rates . the study also shows that smaller language models have the ability to explain HVSB . |