Improving generalization in deep learning is a big deal. The phenomenon is academically interesting either way, but e.g. making sota nets more training data economical seems like a practical result that might be entirely within reach.
i think y'all are both right. grokking is a phenomenon that by definition applies to severely overfit neural networks, which is a very different regime than modern ML - but we might learn something from this that we can use to improve regularization
Comments
That doesn't sound right at all.
Improving generalization in deep learning is a big deal. The phenomenon is academically interesting either way, but e.g. making sota nets more training data economical seems like a practical result that might be entirely within reach.
i think y'all are both right. grokking is a phenomenon that by definition applies to severely overfit neural networks, which is a very different regime than modern ML - but we might learn something from this that we can use to improve regularization
Looks like grokking could give better reasoning and generalization to LLMs, but I'm not sure how practical it would be to overfit a larger LLM
See: https://arxiv.org/abs/2405.15071