Comment on Grokfast: Accelerated Grokking by Amplifying Slow GradientsparentComments−whimsicalism2yi think y'all are both right. grokking is a phenomenon that by definition applies to severely overfit neural networks, which is a very different regime than modern ML - but we might learn something from this that we can use to improve regularization−barfbagginus2yLooks like grokking could give better reasoning and generalization to LLMs, but I'm not sure how practical it would be to overfit a larger LLMSee: https://arxiv.org/abs/2405.15071
Comments
i think y'all are both right. grokking is a phenomenon that by definition applies to severely overfit neural networks, which is a very different regime than modern ML - but we might learn something from this that we can use to improve regularization
Looks like grokking could give better reasoning and generalization to LLMs, but I'm not sure how practical it would be to overfit a larger LLM
See: https://arxiv.org/abs/2405.15071