Papers by Howe Tissue
Entropy Scheduling in Reinforcement Learning for Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | entropy in reinforcement learning functions analogously to the learning rate in LLMs. |
| Approach: | They propose an entropy scheduling system that optimizes different pre-set goals by controlling and scheduling entropicy at each step of the RL process. |
| Outcome: | The proposed method improves AIME2024 from 50.9 to 54.9 within 40 training steps. |