Comment on 4T transistors, one giant chip (Cerebras WSE-3) [video]parentComments−ipsum22yFun fact, I can also train a 24 trillion parameter model on my laptop! Just need to offload weights to the cloud every layer....It's meaningless to say something can train a model that has 24 trillion parameters without specifying the dataset size and time it takes to train.−cryptonector2yI dare say this thing will be many times faster than thrashing your 24T parameters to the cloud.−ipsum22yYeah, but it'll be slower than the equivalent Nvidia GPU cluster.
Comments
Fun fact, I can also train a 24 trillion parameter model on my laptop! Just need to offload weights to the cloud every layer.
...
It's meaningless to say something can train a model that has 24 trillion parameters without specifying the dataset size and time it takes to train.
I dare say this thing will be many times faster than thrashing your 24T parameters to the cloud.
Yeah, but it'll be slower than the equivalent Nvidia GPU cluster.