The claim (which is not uncontested, I should add) is that doing so repeatedly inevitably produces model collapse. Even if that is true, however, you can still derive benefit from using larger models to generate large amounts of synthetic training data for smaller models. Most LLaMA finetunes out there are trained on GPT-4 output, for example.
Comments
I'm so confused.
Isn't "training LLMs on LLM output" the very definition of "model collapse" or "model poisoning"?
The claim (which is not uncontested, I should add) is that doing so repeatedly inevitably produces model collapse. Even if that is true, however, you can still derive benefit from using larger models to generate large amounts of synthetic training data for smaller models. Most LLaMA finetunes out there are trained on GPT-4 output, for example.