Skip to content

Comment on New attention mechanisms that outperform standard multi-head attentionparent

Comments

I am not sure why this article was posted here on hackernews.

New is where progress comes from, so new is interesting. New is why we come here, and the first three letters of News.

here is very little theoretic about transformer-style architectures.

Only way to fix that is with new.

Fundamentally, the proof is in the pudding, not in "mathematical comparisons"

"Can it scale" is something only someone with money can answer. It can be tested, but only if it's known. Now new is better known.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.