Papers by Ah Seo

    1 papers
    Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback (2024.findings-emnlp)

    Copied to clipboard

    Challenge: Existing methods for large language models rely on binary labels that fail to capture the subtle differences in relative quality between pairs.
    Approach: They propose a method that incorporates relative quality margins into optimization to improve LLM policies and reward models.
    Outcome: The proposed approach outperforms baseline methods on popular benchmarks including MT-bench and RewardBench.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations