I'm going to share a, probably, controversial opinion. That opinion is: I can't stand an outage title like "Websites and APIs on Render are unavailable due to Cloudflare network errors". Its passing blame. I run an app or two on Render. I don't pay Cloudflare; I pay Render. Take responsibility for the infrastructure decisions that you make, for your customers; don't pass blame to your infrastructure providers.
We take full responsibility for the infrastructure choices that led to this outage. As the peer comment said, it's helpful to overshare in these situations.
We know developers don't actually care who's at fault and will move off of Render if we're down, period. Even before the incident, we'd started working on a project to eliminate the SPOF with Cloudflare, and now it's only a matter of time before we ship it.
I get that, and the update is much appreciated. I don't mean to insinuate that this was the intention behind why that language was chosen; its just the sentiment that the language conveys, and that's why I'm not a fan of it.
The stance that I take is; its a fine line between Oversharing and Passing Blame in outages like this, and while I'm happy that a line like that when shared by Render means it was just oversharing (I love your product!), its easy to see how a line like that when shared by a less admirable company could be seen as "Nah man, its not on us, we didn't do anything wrong." A critical difference being; if Cloudflare was the cause, how are we working toward avoiding this cause in the future; which leads nicely to where pointing at Cloudflare (or any upstream provider) generally feels more agreeable; the retro.
To be clear; I have no intention of leaving Render, even if y'all weren't planning to alleviate this SPOF. I fully grok the difficult engineering required to nuke SPOFs like Cloudflare or AWS; and a bit of downtime here and there is a price I'm fine with paying.
Sure, but there's a time and a place. Outages involve high tensions and fog of war; and you said it yourself, you're already ready to blame them for not having backup generators in this hypothetical example. The midst of an outage is not the time to start casting blame, on people, organizations, processes, providers, whatever. Outages are the time to fix; retros are the time to blame (within productive reason, of course).
If you ran it outside of Render, would you be using a CDN service or building your own?
The bigger issue you're alluding to is that of supply-chain reliability in SAAS products: when AWS goes down, multiple other (seemingly unrelated) services go down. But saying its the downstream service's fault is pointless, because if you were to do it yourself you'd be using the same upstream provider, and be dealing with their outage yourself.
In that example, Slack as a bigger of AWS would have a much bigger say, and a more direct line to AWS engineers, than you would.
Right, and I think there's an interesting transitive correlation here: As a customer of Render, while Render was down because of Cloudflare; is it appropriate for me to post on our outage page: "Service interruption due to issues at Render"? "Service interruption due to issues at Cloudflare"? What does Cloudflare post on their page? (Well, they may actually post "due to a busted AC unit in our Seattle data center" which, you know, at that point we've hit bedrock so maybe that's valuable, but)
Its turtles all the way down, and in the midst of an outage I totally empathize with the off-the-cuff thinking that oversharing is better than undersharing, but after the fog of war clears you can even retro language like that and come to a different conclusion. What value do my customers, even if they're highly technical, gain by knowing its Render's fault that MyCoolService was down? Are they going to go open support tickets with Render? I'd bet Render very reasonably wouldn't appreciate that, and they're not going to have a better trunk to their support than I do.
Comments
I'm going to share a, probably, controversial opinion. That opinion is: I can't stand an outage title like "Websites and APIs on Render are unavailable due to Cloudflare network errors". Its passing blame. I run an app or two on Render. I don't pay Cloudflare; I pay Render. Take responsibility for the infrastructure decisions that you make, for your customers; don't pass blame to your infrastructure providers.
We take full responsibility for the infrastructure choices that led to this outage. As the peer comment said, it's helpful to overshare in these situations.
We know developers don't actually care who's at fault and will move off of Render if we're down, period. Even before the incident, we'd started working on a project to eliminate the SPOF with Cloudflare, and now it's only a matter of time before we ship it.
I get that, and the update is much appreciated. I don't mean to insinuate that this was the intention behind why that language was chosen; its just the sentiment that the language conveys, and that's why I'm not a fan of it.
The stance that I take is; its a fine line between Oversharing and Passing Blame in outages like this, and while I'm happy that a line like that when shared by Render means it was just oversharing (I love your product!), its easy to see how a line like that when shared by a less admirable company could be seen as "Nah man, its not on us, we didn't do anything wrong." A critical difference being; if Cloudflare was the cause, how are we working toward avoiding this cause in the future; which leads nicely to where pointing at Cloudflare (or any upstream provider) generally feels more agreeable; the retro.
To be clear; I have no intention of leaving Render, even if y'all weren't planning to alleviate this SPOF. I fully grok the difficult engineering required to nuke SPOFs like Cloudflare or AWS; and a bit of downtime here and there is a price I'm fine with paying.
Heard, loud and clear. Thanks for the support.
If Renders data center was down due to a city wide power outage I would still like to know because it’s the root cause.
I would still blame them for not having back up generators though.
However a failure to plan for emergencies is different from other kinds of failure.
Sure, but there's a time and a place. Outages involve high tensions and fog of war; and you said it yourself, you're already ready to blame them for not having backup generators in this hypothetical example. The midst of an outage is not the time to start casting blame, on people, organizations, processes, providers, whatever. Outages are the time to fix; retros are the time to blame (within productive reason, of course).
If you ran it outside of Render, would you be using a CDN service or building your own?
The bigger issue you're alluding to is that of supply-chain reliability in SAAS products: when AWS goes down, multiple other (seemingly unrelated) services go down. But saying its the downstream service's fault is pointless, because if you were to do it yourself you'd be using the same upstream provider, and be dealing with their outage yourself.
In that example, Slack as a bigger of AWS would have a much bigger say, and a more direct line to AWS engineers, than you would.
Right, and I think there's an interesting transitive correlation here: As a customer of Render, while Render was down because of Cloudflare; is it appropriate for me to post on our outage page: "Service interruption due to issues at Render"? "Service interruption due to issues at Cloudflare"? What does Cloudflare post on their page? (Well, they may actually post "due to a busted AC unit in our Seattle data center" which, you know, at that point we've hit bedrock so maybe that's valuable, but)
Its turtles all the way down, and in the midst of an outage I totally empathize with the off-the-cuff thinking that oversharing is better than undersharing, but after the fog of war clears you can even retro language like that and come to a different conclusion. What value do my customers, even if they're highly technical, gain by knowing its Render's fault that MyCoolService was down? Are they going to go open support tickets with Render? I'd bet Render very reasonably wouldn't appreciate that, and they're not going to have a better trunk to their support than I do.