Papers by Naipeng Chao
Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used in role-play scenarios, but their safety implications remain under-characterized. |
| Approach: | They propose a diagnostic benchmark for role-play jailbreaks based on Bandura’s Moral Disengagement theory and propose 'MD-Trace' based defense that reduces attack success while maintaining Role Fidelity. |
| Outcome: | The proposed framework improves safety behavior for benign personas while increasing unsafe compliance for malicious ones. |
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods to learn internal world models rely on one-step supervision . however, standard MTP suffers from structural hallucinations . |
| Approach: | They propose a method which anchors predictions to ground-truth hidden state trajectories. |
| Outcome: | The proposed method bridges the gap between discrete tokens and continuous state representations, reducing structural hallucinations, and improving robustness to perturbations. |
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing research focuses on character-level settings and static evaluation formats fail to capture the complexity of everyday social interactions. |
| Approach: | They propose a dynamic simulation framework for evaluating and improving persona-level role-playing in large language models (LLMs). |
| Outcome: | The proposed framework leverages user-generated social content to construct a nuanced persona bank and elicits multi-turn, context-rich interactions within simulated social environments. |