Papers by Mengqi Yuan
Direct Multi-Turn Preference Optimization for Language Agents (2024.emnlp-main)
Copied to clipboard
| Challenge: | Extensive experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the DMPO loss function. |
| Approach: | They propose a novel loss function for multi-turn agent tasks that replaces the policy constraint with the state-action occupancy measure constraint and adds length normalization to the Bradley-Terry model. |
| Outcome: | Experiments on three multi-turn agent task datasets confirm the effectiveness and superiority of the proposed loss function. |