Skip to content

Comment on Scaling Laws, Carefullyparent

Comments

Isn't that one of the reasons why KL-divergence is used, at least in DPO/RL for LLM? Otherwise the model can effectively cheat and mode collapse. For pre-training against a 1-hot label the KL-divergence should be equivalent to cross-entropy anyway.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.