Papers by Kaiqu Liang

1 papers
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation (2026.findings-acl)

Copied to clipboard

Challenge: Reinforcement Learning from Hindsight Simulation (RLHF) can cause severe misalignment in generative AI, but it is not a universal method for fine-tuning large language models.
Approach: They propose a method that uses evaluator feedback to decouple alignment signal from potentially compromised predictions.
Outcome: The proposed method significantly outperforms RLHF in comparisons with baselines and human evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations