Skip to content

Mamba: New SSM arch with linear-time scaling that outperforms Transformers

github.com/state-spaces
6 pointsxenova2 comments
On HN

Comments

This is promising as the future of modeling. Outperforming transformers doesn’t catch my eye because there is so much variation in performance based on size, training data, training method etc.

However the greater (5x) inference bandwidth makes this super appealing especially for democratizing AI and enabling the GPU poor. This could very well be a watershed moment for SSMs, similar to how Transformers boasted improvements to training and inference speed in 2019.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.