Papers by Mengqi Yuan

    1 papers
    Direct Multi-Turn Preference Optimization for Language Agents (2024.emnlp-main)

    Copied to clipboard

    Challenge: Extensive experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the DMPO loss function.
    Approach: They propose a novel loss function for multi-turn agent tasks that replaces the policy constraint with the state-action occupancy measure constraint and adds length normalization to the Bradley-Terry model.
    Outcome: Experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the proposed loss function.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations