Papers by Zhenyu Weng
Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing questionnaire-based evaluation methods exhibit limited stability and offer little explainability, as their results are sensitive to minor variations in prompt phrasing or role-play configurations. |
| Approach: | They propose an internal-activation-based approach for stable and explainable personality trait evaluation in Large Language Models by interpolating a persona vector associated with a target personality trait from the model's internal activations. |
| Outcome: | The proposed approach yields significantly more stable personality trait evaluations than existing methods, even under questionnaire and role-play variants. |