Papers by Lior Belenki
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods to optimize language model pre-training data mixtures are difficult due to the complexity of the data mixture. |
| Approach: | They propose a method to optimize language model pre-training data mixtures by approximating cross-entropy loss via a Mixture of Data Experts (MDE). |
| Outcome: | The proposed method improves performance on a slimPajama dataset with a mixture of data experts. |