Skip to content

Comment on New attention mechanisms that outperform standard multi-head attention

Comments

The models tested are extremely small, a few thousand parameters and the performance is of course not great, I don't think we can extrapolate much from this. I don't understand why they chose such small models when you can train much larger ones for free on Colab or Kaggle if you really need it.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.