Papers by Mingye Zhu

    2 papers
    LIRE: listwise reward enhancement for preference alignment (2024.findings-acl)

    Copied to clipboard

    Challenge: prevailing approaches to preference alignment focus on pairwise comparisons, with limited exploration into multi-response scenarios.
    Approach: They propose a listwise reward enhancement approach that integrates offline rewards of multiple responses into a streamlined listwise framework.
    Outcome: The proposed approach outperforms existing methods on dialogue and summarization tasks with good transferability to out-of-distribution data.
    FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization (2024.emnlp-main)

    Copied to clipboard

    Challenge: Recent advances in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values.
    Approach: They propose a constrained optimization approach to detect and mitigate update regression with focal attention.
    Outcome: The proposed approach detects and mitigates update regression with focal attention while maintaining excellent overall performance.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations