Sorry if the joke didn't land; I have heard a lot of different numbers for the size of US labs' models, but never seen any of them substantiated, so I think you're likely to just get more rumours in answer to this question.
My personal take, with no sources: 100T sounds excessively high given they need to be able to actually serve these things on commercially available hardware. I would guess they are in the same order of magnitude as the Chinese frontier models. It's possible their edge is in RL training methods, training-time compute, and access to data (e.g. from customers' CC/Codex sessions), not in model size.
Allright! Well, the rumors I got are around 100T for those frontier models: reading stuff here and there on internet, on discussion forums with guys pretending running AI models, etc.
This is so hard to sort the true from the false nowadays.
If LLM becomes that good at coding, I'll have to run an open weight one locally.
Comments
Sorry if the joke didn't land; I have heard a lot of different numbers for the size of US labs' models, but never seen any of them substantiated, so I think you're likely to just get more rumours in answer to this question.
My personal take, with no sources: 100T sounds excessively high given they need to be able to actually serve these things on commercially available hardware. I would guess they are in the same order of magnitude as the Chinese frontier models. It's possible their edge is in RL training methods, training-time compute, and access to data (e.g. from customers' CC/Codex sessions), not in model size.
Allright! Well, the rumors I got are around 100T for those frontier models: reading stuff here and there on internet, on discussion forums with guys pretending running AI models, etc.
This is so hard to sort the true from the false nowadays.
If LLM becomes that good at coding, I'll have to run an open weight one locally.