Papers by Roman Plaud
Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies show that large language models can learn to achieve long-term goals when fine-tuned with reinforcement learning. |
| Approach: | They propose a modification to the Kullback-Leibler penalty to favor exploration on critical tokens . they show how varying degrees of pre-training influence exploration . |
| Outcome: | The proposed model favors exploration on critical tokens, increasing the efficiency of the RL fine-tuning stage. |