Papers by Chris Cremer
Back to Basics: Revisiting REINFORCE-Style Optimization for Learning from Human Feedback in LLMs (2024.acl-long)
Copied to clipboard
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, Sara Hooker
| Challenge: | Proximal Policy Optimization (PPO) is used for RLHF but requires high computational cost and sensitive hyperparameter tuning. |
| Approach: | They propose to use Proximal Policy Optimization to align large language models to human preferences. |
| Outcome: | The proposed method preserves and even increases performance while preserving the motivational principles that led to the development of PPO. |
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion (2024.emnlp-main)
Copied to clipboard
Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Azar, Olivier Pietquin, Matthieu Geist
| Challenge: | Reinforcement Learning (RL) is a method used to fine tune Large Language Models (LLMs) using a reward model trained from preference data to better align with human judgment. |
| Approach: | They propose a Reinforcement Learning (RL) algorithm that can estimate the optimal policy even from off-policy data. |
| Outcome: | The proposed algorithm can estimate the optimal policy even from off-policy data. |