Skip to content

Comment on New attention mechanisms that outperform standard multi-head attentionparent

Comments

Why, then, we barely see any non-transformer production-ready LLM these days?

Because having a 5% better non-transformer model doesn't help you if as a result you can't use the 10% improvements people publish that only apply to transformers. Very quickly you'll be 5% worse than those who stuck with transformers, and have wasted a ton of time and money.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.