Comment on New attention mechanisms that outperform standard multi-head attentionComments−lalaland11252yThe Transformer models, used in this experiment, all have a single attention layer with model dimension and context length 32.I think we are going to need to see more experiments here, especially because the theoretical motivations here are weak−ttul2yI’m certainly not a domain expert, but one thing I have read repeatedly about Transformers is that not all tricks scale the same.
Comments
I think we are going to need to see more experiments here, especially because the theoretical motivations here are weak
I’m certainly not a domain expert, but one thing I have read repeatedly about Transformers is that not all tricks scale the same.