Skip to content

Comment on Grokfast: Accelerated Grokking by Amplifying Slow Gradientsparent

Comments

grokking doesn't and will not have practical uses, imo - it is just an experiment that revealed cool things that we mostly already suspected about implicit regularization

however, techniques we learn from grokking about implicit regularization might be helpful for the training regimes we actually use

grokking doesn't and will not have practical uses, imo

I'm not so sure. Reasoning is the next big hurdle, and grokking and parametric memory seem very effective here.

[1] https://arxiv.org/abs/2405.15071

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.