Skip to content

Comment on Nemotron-4-340Bparent

Comments

The claim (which is not uncontested, I should add) is that doing so repeatedly inevitably produces model collapse. Even if that is true, however, you can still derive benefit from using larger models to generate large amounts of synthetic training data for smaller models. Most LLaMA finetunes out there are trained on GPT-4 output, for example.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.