Papers with DA-ILQL
GOODLIAR: A Reinforcement Learning-Based Deceptive Agent for Disrupting LLM Beliefs on Foundational Principles (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances indicate that LLMs exhibit increasingly complex reasoning abilities . |
| Approach: | They propose a reinforcement learning framework that generates deceptive contexts to rewrite an LLM’s core axiomatic beliefs. |
| Outcome: | The proposed framework induces persistent belief shifts rather than one-off policy breaches. |