Papers with WildReward

    1 papers
    WildReward: Learning Reward Models from In-the-Wild Human Interactions (2026.acl-long)

    Copied to clipboard

    Challenge: Prior work focused on collecting preference pairs, requiring substantial annotation efforts.
    Approach: They propose a pipeline to extract reliable human feedback from in-the-wild interactions . they propose to use WildChat as an interaction source to train the model .
    Outcome: The proposed model achieves comparable or even superior performance compared to conventional models with improved calibration and cross-sample consistency.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations