Papers by Daoyi Dong

    1 papers
    Select Before Use: On the Importance of Reference Model Selection in Preference Alignment (2026.acl-long)

    Copied to clipboard

    Challenge: Supervised Fine-Tuning (SFT) is used as the initialization and reference model for subsequent preference alignment.
    Approach: They propose to use RewardRank to estimate initial implicit alignment between reference model and preference objective to ensure LLMs generate safe, helpful, and instruction-aligned content.
    Outcome: Empirical evidence shows that using the selected model as reference can gain up to 67.6% relative increase on length-controlled win rate compared to baselines.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations