Papers by Taneesh Gupta

    1 papers
    CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling (2025.findings-acl)

    Copied to clipboard

    Challenge: Reward modeling in large language models is susceptible to reward hacking . flawed reward signals often lead to outputs that optimize for spurious correlates .
    Approach: They propose a new approach that generates dynamic, context-relevant criteria to ground the reward model prior to producing reward scores.
    Outcome: The proposed approach generates dynamic, context-relevant criteria to ground the model prior to producing reward scores.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations