Skip to content

Comment on Grokfast: Accelerated Grokking by Amplifying Slow Gradientsparent

Comments

They are trying to innovate in the idea space and are probably quite compute constrained.

Training a GPT-2 sized model costs ~$20 nowadays in respect to compute: https://github.com/karpathy/llm.c/discussions/481

Baseline time to grok something looks to be around 1000x normal training time so make that $20k per attempt. Probably takes a while too. Their headline number (50x faster than baseline, $400) looks pretty doable if you can make grokking happen reliably at that speed.

$20 per attempt. A paper typically comes after trying hundreds of things. That said, the final version of your idea could certainly try it.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.