Skip to content

Comment on New attention mechanisms that outperform standard multi-head attention

Comments

The Transformer models, used in this experiment, all have a single attention layer with model dimension and context length 32.

I think we are going to need to see more experiments here, especially because the theoretical motivations here are weak

I’m certainly not a domain expert, but one thing I have read repeatedly about Transformers is that not all tricks scale the same.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.