Papers by Kaiqu Liang
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation (2026.findings-acl)
Copied to clipboard
| Challenge: | Reinforcement Learning from Hindsight Simulation (RLHF) can cause severe misalignment in generative AI, but it is not a universal method for fine-tuning large language models. |
| Approach: | They propose a method that uses evaluator feedback to decouple alignment signal from potentially compromised predictions. |
| Outcome: | The proposed method significantly outperforms RLHF in comparisons with baselines and human evaluations. |