I have a suspicion that this technique will prove most valuable for market oriented data sets (like price related time series), where there isn't necessarily that much massive data scale compared to text corpora, and where there are very tight limits on the amount of training data because you only want to include recent data to reduce the chances of market regime changes. This approach seems to shine when you don't quite have enough training data to completely map out the general case, but if you train for long enough naively, you can get lucky and fall into it.
Comments
I have a suspicion that this technique will prove most valuable for market oriented data sets (like price related time series), where there isn't necessarily that much massive data scale compared to text corpora, and where there are very tight limits on the amount of training data because you only want to include recent data to reduce the chances of market regime changes. This approach seems to shine when you don't quite have enough training data to completely map out the general case, but if you train for long enough naively, you can get lucky and fall into it.