Papers by Mehwish Nasim
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on personality steering in large language models rely on injecting trait-specific steering vectors into the residual stream to control the strength of trait expression. |
| Approach: | They examine the geometric relationships between Big Five personality steering directions by applying geometric conditioning schemes to their steering vectors. |
| Outcome: | The proposed model can be used to steer personality traits in large language models. |
Can LLM Agents Maintain a Persona in Discourse? (2025.emnlp-main)
Copied to clipboard
Pranav Bhandari, Nicolas Fay, Michael J Wise, Amitava Datta, Stephanie Meek, Usman Naseem, Mehwish Nasim
| Challenge: | Large language models are often subjected to context-shifting behaviour, resulting in a lack of consistent and interpretable personality-aligned interactions. |
| Approach: | They propose to use two conversation agents to generate a discourse with an assigned personality from the OCEAN framework and then use multiple judge agents to infer original traits. |
| Outcome: | The proposed model is based on two conversation agents with a personality assigned from the OCEAN framework and then multiple judge agents to infer the original traits assigned. |
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Personality-aware LLMs exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. |
| Approach: | They propose a pipeline that extracts hidden state activations from transformer layers using the Big Five Personality Traits framework. |
| Outcome: | The proposed model extracts hidden state activations from transformer layers using the Big Five personality traits (Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism) |
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection (2026.acl-long)
Copied to clipboard
| Challenge: | censored models outperform uncensoreed counterparts in accuracy and robustness, achieving 69.0% accuracy versus 64.1% strict accuracy. |
| Approach: | They examine how large language models with minimal safety alignment compare with more heavily aligned counterparts when deployed using political personas. |
| Outcome: | The proposed model outperforms uncensored models in accuracy and robustness, while uncensors are more malleable to ideological framing. |