Papers by Sanjeevan Selvaganapathy
Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Personality-aware LLMs exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. |
| Approach: | They propose a pipeline that extracts hidden state activations from transformer layers using the Big Five Personality Traits framework. |
| Outcome: | The proposed model extracts hidden state activations from transformer layers using the Big Five personality traits (Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism) |
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection (2026.acl-long)
Copied to clipboard
| Challenge: | censored models outperform uncensoreed counterparts in accuracy and robustness, achieving 69.0% accuracy versus 64.1% strict accuracy. |
| Approach: | They examine how large language models with minimal safety alignment compare with more heavily aligned counterparts when deployed using political personas. |
| Outcome: | The proposed model outperforms uncensored models in accuracy and robustness, while uncensors are more malleable to ideological framing. |