Papers by Jiayong Wan
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing defense methods are insufficient to address in-context reward hacking (ICRH), where LLMs iteratively optimize their behavior to maximize proxy objectives, resulting in harmful side effects. |
| Approach: | They propose a framework that reduces in-context reward hacking (ICRH) through repeated interactions with the environment. |
| Outcome: | The proposed framework reduces ICRH without model fine-tuning while maintaining task performance. |