Comment on New attention mechanisms that outperform standard multi-head attentionparentComments−janalsncm2ySometimes it doesn’t need to. You might have a problem that isn’t web scale and where transfer learning is hard. We also need techniques for small datasets even if they are slower to train or are outperformed after 5 billion tokens.
Comments
Sometimes it doesn’t need to. You might have a problem that isn’t web scale and where transfer learning is hard. We also need techniques for small datasets even if they are slower to train or are outperformed after 5 billion tokens.