Dario Amodei has apparently recently suggested that Anthropic might become only only private AI company in the entire world, which obviously it won't.
There is competition everywhere, and it is intensifying and catching up, not fading away. Open weight models are becoming more common, both within the US as well as elsewhere. Treasury secretary Scott Bessent just praised Meta's open weight models.
There is demand for AI at all different price points, and as all models at all price points become more capable, it seems that increasingly developers are seeing the most expensive ones as specialized tools, not daily drivers.
Compute/memory may be constrained for a few years until production capacity catches up, but this does not mean that demand for cheaper and open weight models will go away, else it would already be happening. Anthropic would like to sell an expensive Ferrari to everyone on the planet, but 99.99% of those people have no need for anything more than a Yugo.
No - but we already have cheap LLMs priced way below frontier models. This is not the housing market. There will always be someone willing to take a lower profit margin for a slice of the pie, and of course smaller models are cheaper to serve so can afford to be cheaper.
DeepSeek recently said that their super-low pricing let's them recoup the cost of the hardware it runs on in 10 months, so there is evidentially plenty of profit to be had over a projected 3+ year lifespan of a "GPU".
Some in the AI industry, or breathing the same air (Dwarkesh) project that limited GPUs will only be used to serve the most expensive models with the highest profit margins, but it is just not what we are seeing. If the only LLMs available were ones at Opus/Fable price points then the GPU scarcity would disappear since the demand at that price is just not there. It's remarkably like trying to fill all the seats on a plane - you can fill a few at 1st class prices, but most of the plane better be coach if you want to sell all the seats.
For a GPU, "selling all the seats", keeping it busy 24x7, is critical to profitability since the primary cost to serving is the GPU which has a limited lifespan.
There's a huge range of cost-per-task variation across models, and the capability of the smaller cheaper models keeps increasing.
For example, here we have Fable 5 at $3.14/task vs Kimi K3 at $0.84/task, with very little difference between them in coding capability (and this isn't even a coding/agentic fine tune of Kimi).
We now have models like Qwen 3.8 27B, small enough to run locally, with coding capability similar to Opus 4.5 based on challenging tasks like the Anthropic Kernel challenge.
I think we are rapidly getting to the "good enough" stage of LLMs, just like we did long ago with PCs. A cheap PC/LLM is all you need for 99.9% of normal use cases. Maybe nothing can touch whatever latest greatest models Anthropic and OpenAI have when it comes to solving Erdos problems, but most developers are working on problems more like the Anthropic Kernel challenge in complexity (or in fact typically way simpler ones).
Look at housing, we've been waiting 20 years for "production capacity to catch up".
A large portion of the US construction capacity is now engaged in building data centers, priorities you know. Who knows what else will be a favorite tomorrow but it's unlikely to be properly built housing.
IMO this also clearly shows that the housing shortages happening in so many countries could have been fixed all along if those hording most resources wanted to, but why would they? The shortages are heavily tipping the scales in their favor.
Comments
Dario Amodei has apparently recently suggested that Anthropic might become only only private AI company in the entire world, which obviously it won't.
There is competition everywhere, and it is intensifying and catching up, not fading away. Open weight models are becoming more common, both within the US as well as elsewhere. Treasury secretary Scott Bessent just praised Meta's open weight models.
There is demand for AI at all different price points, and as all models at all price points become more capable, it seems that increasingly developers are seeing the most expensive ones as specialized tools, not daily drivers.
Compute/memory may be constrained for a few years until production capacity catches up, but this does not mean that demand for cheaper and open weight models will go away, else it would already be happening. Anthropic would like to sell an expensive Ferrari to everyone on the planet, but 99.99% of those people have no need for anything more than a Yugo.
Demand for cheap stuff doesn't manifest cheap stuff. Look at housing, we've been waiting 20 years for "production capacity to catch up".
No - but we already have cheap LLMs priced way below frontier models. This is not the housing market. There will always be someone willing to take a lower profit margin for a slice of the pie, and of course smaller models are cheaper to serve so can afford to be cheaper.
DeepSeek recently said that their super-low pricing let's them recoup the cost of the hardware it runs on in 10 months, so there is evidentially plenty of profit to be had over a projected 3+ year lifespan of a "GPU".
Some in the AI industry, or breathing the same air (Dwarkesh) project that limited GPUs will only be used to serve the most expensive models with the highest profit margins, but it is just not what we are seeing. If the only LLMs available were ones at Opus/Fable price points then the GPU scarcity would disappear since the demand at that price is just not there. It's remarkably like trying to fill all the seats on a plane - you can fill a few at 1st class prices, but most of the plane better be coach if you want to sell all the seats.
For a GPU, "selling all the seats", keeping it busy 24x7, is critical to profitability since the primary cost to serving is the GPU which has a limited lifespan.
As we have seen, those cheap LLMs are cheap because their volume is so low, not because they aren't greedy or discovered some kind of efficiency hack.
There is a clear trend of popularity and price.
There's a huge range of cost-per-task variation across models, and the capability of the smaller cheaper models keeps increasing.
For example, here we have Fable 5 at $3.14/task vs Kimi K3 at $0.84/task, with very little difference between them in coding capability (and this isn't even a coding/agentic fine tune of Kimi).
https://artificialanalysis.ai/models
We now have models like Qwen 3.8 27B, small enough to run locally, with coding capability similar to Opus 4.5 based on challenging tasks like the Anthropic Kernel challenge.
I think we are rapidly getting to the "good enough" stage of LLMs, just like we did long ago with PCs. A cheap PC/LLM is all you need for 99.9% of normal use cases. Maybe nothing can touch whatever latest greatest models Anthropic and OpenAI have when it comes to solving Erdos problems, but most developers are working on problems more like the Anthropic Kernel challenge in complexity (or in fact typically way simpler ones).
A large portion of the US construction capacity is now engaged in building data centers, priorities you know. Who knows what else will be a favorite tomorrow but it's unlikely to be properly built housing.
IMO this also clearly shows that the housing shortages happening in so many countries could have been fixed all along if those hording most resources wanted to, but why would they? The shortages are heavily tipping the scales in their favor.