Papers by Qianben Chen

    1 papers
    DUAL RM: Beyond Rule-based Preference Reward Modeling via Meta-Reward (2026.acl-long)

    Copied to clipboard

    Challenge: Existing preference-based reward modeling methods face a recursive dependency where each verifier requires a meta-verifier, leading to continuous and costly dependence on human annotation.
    Approach: They propose a dual RM that couples discriminative and generative reward models under a non-parametric meta-reward.
    Outcome: The proposed model achieves strong performance across major preference benchmarks and even when trained exclusively on language modality, it exhibits robust cross-modal transfer on Omni-RewardBench.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations