Papers by Nima Chitsazan
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning (2025.acl-long)
Copied to clipboard
| Challenge: | Increasing language model size improves cross-entropy loss with power-law behaviour, but scaling laws do not explain how scaling improves loss. |
| Approach: | They find that language models undergo loss deceleration early in training . they attribute loss deceleration to a type of degenerate training dynamics we call zero-sum learning . |
| Outcome: | The proposed scaling improves loss on language models, but degrades loss in other subsets, resulting in bottlenecks. |