Skip to content

Comment on Grokfast: Accelerated Grokking by Amplifying Slow Gradients

Comments

I have a suspicion that this technique will prove most valuable for market oriented data sets (like price related time series), where there isn't necessarily that much massive data scale compared to text corpora, and where there are very tight limits on the amount of training data because you only want to include recent data to reduce the chances of market regime changes. This approach seems to shine when you don't quite have enough training data to completely map out the general case, but if you train for long enough naively, you can get lucky and fall into it.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.