Papers with MetaPO

    1 papers
    Learning Temporally-Aware Sample Weights for Preference Optimization (2026.findings-acl)

    Copied to clipboard

    Challenge: Existing methods for preference optimization rely on static functions of instantaneous model states and ignore temporal learning dynamics.
    Approach: They propose a framework that meta-learns adaptive weights using three temporal features: reward margin evolution, learning volatility, and reference deviation.
    Outcome: The proposed framework achieves statistically significant improvements over baselines on models ranging from 7B to 70B parameters.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations