Papers by Mathieu Rita
Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning. |
| Approach: | They propose a reinforcement learning approach that leverages human demonstrations and a reward model to recalibrate the reward objective. |
| Outcome: | The proposed approach achieves comparable performance to carefully tuned baselines while mitigating ROO in three RL language tasks. |
On the Correspondence between Compositionality and Imitation in Emergent Neural Communication (2023.findings-acl)
Copied to clipboard
| Challenge: | a study examining compositionality and imitation learning in a Lewis game demonstrates that it is difficult to imitate compositional languages. |
| Approach: | They explore the link between compositionality and imitation in a Lewis game . they show that the learning algorithm used to imitate is crucial . |
| Outcome: | The proposed model improves compositionality and imitation in a Lewis game . the study shows that compositional languages are easier to imitate . |