Papers by Rémy Portelas
Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used in the creation of online content, creating feedback loops as future generations of models will be trained on this synthetic data. |
| Approach: | They propose to use large language models to create feedback loops as future models are trained on this data. |
| Outcome: | The proposed model collapse effects are found to be detrimental to the results of recursive training on human datasets. |