It is a reliability problem - because if the API ends up rejecting 30% of the incoming queries, no one cares if it's 500 Internal Server Error or 529 Overloaded.
Having your infrastructure at the knife's edge of load to capacity also means you have no redundancy when something fails.
If this incident had happened during the day I could buy that argument. But it is happening in the dead of the night. Are there really that much people scheduling claude runs over-night, taking more capacity than day-time? (assume most claude users are west coast programmers)
My other guess is they are scheduling training runs and their capacity isolation is bad.
Only if you're incapable of throttling to assure quality of service. Are you saying they have so much demand that even their load balancers were overloaded?
Throttling helps to prevent a 30% failure rate (with a quick recovery time) turning into a 100% failure rate (with practically infinite recovery time), but some users are still going to see failures, and the news is still going to be "Anthropic is down again" because some users are seeing an outage.
I'm not really sure how they can avoid getting in the news with "Anthropic is down", TBH. Although, at least by throttling they might get "Anthropic is down for some requests" and not "Anthropic is down for all requests".
There is no point having this discussion. Some people have had their brains genuinely broken by Anthropic and will defend them to the last breath, despite having no clue what they are talking about.
Comments
Because they can’t keep up with the demand?
Struggling with demands is a performance problem, not a reliability problem.
It is a reliability problem - because if the API ends up rejecting 30% of the incoming queries, no one cares if it's 500 Internal Server Error or 529 Overloaded.
Having your infrastructure at the knife's edge of load to capacity also means you have no redundancy when something fails.
If this incident had happened during the day I could buy that argument. But it is happening in the dead of the night. Are there really that much people scheduling claude runs over-night, taking more capacity than day-time? (assume most claude users are west coast programmers)
My other guess is they are scheduling training runs and their capacity isolation is bad.
The world is more than just the Americas.
It wasn’t night for all of Asia or Europe?
overly high demand can often put a system in an unstable state
Only if you're incapable of throttling to assure quality of service. Are you saying they have so much demand that even their load balancers were overloaded?
Throttling helps to prevent a 30% failure rate (with a quick recovery time) turning into a 100% failure rate (with practically infinite recovery time), but some users are still going to see failures, and the news is still going to be "Anthropic is down again" because some users are seeing an outage.
I'm not really sure how they can avoid getting in the news with "Anthropic is down", TBH. Although, at least by throttling they might get "Anthropic is down for some requests" and not "Anthropic is down for all requests".
There is no point having this discussion. Some people have had their brains genuinely broken by Anthropic and will defend them to the last breath, despite having no clue what they are talking about.