Papers by Fanzhi Zeng

    1 papers
    Reward Generalization in RLHF: A Topological Perspective (2025.findings-acl)

    Copied to clipboard

    Challenge: Existing alignment methods share a common topology of information flow, but their alternatives have not been thoroughly explored.
    Approach: They propose a theory of reward generalization in reinforcement learning from human feedback . they propose induced Bayesian networks to model the impact of dataset topologies on reward generalisation .
    Outcome: The proposed method achieves an average win rate of 65% on three NLP tasks.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations