Papers with off-policy
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout (N19-1)
Copied to clipboard
| Challenge: | Existing approaches perform significantly worse in unseen environments compared to seen ones. |
| Approach: | They propose to use a ‘environmental dropout’ method to generate unseen triplets to generate new paths and instructions to generalize the agent. |
| Outcome: | The proposed agent outperforms the state-of-the-art approaches on the private unseen test set and is ranked top on the leaderboard. |
WPO: Enhancing RLHF with Weighted Preference Optimization (2024.emnlp-main)
Copied to clipboard
Wenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu
| Challenge: | Off-policy preference optimization suffers from a distributional gap between the policy used for data collection and the target policy, leading to suboptimal optimization. |
| Approach: | They propose a method to simulate on-policy learning with off-police preference data. |
| Outcome: | The proposed method outperforms Direct Preference Optimization (DPO) by up to 5.6% on Alpaca Eval 2 and MT-bench. |
Optimizing Reasoning for Text-to-SQL with Execution Feedback (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models excel in many reasoning tasks, but their ability to leverage Chain-of-Thought (CoT) reasoning remains underexplored. |
| Approach: | They propose a framework that iteratively optimizes open-source LLMs by combining CoT reasoning with off-policy and on-poly DPO, relying solely on execution accuracy as feedback. |
| Outcome: | The proposed framework improves execution accuracy on BIRD and Spider datasets. |