Skip to content

Comment on 4T transistors, one giant chip (Cerebras WSE-3) [video]parent

Comments

Fun fact, I can also train a 24 trillion parameter model on my laptop! Just need to offload weights to the cloud every layer.

...

It's meaningless to say something can train a model that has 24 trillion parameters without specifying the dataset size and time it takes to train.

I dare say this thing will be many times faster than thrashing your 24T parameters to the cloud.

Yeah, but it'll be slower than the equivalent Nvidia GPU cluster.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.