Skip to content

Comment on What Powers Instagram: Hundreds of Instances, Dozens of Technologies

Comments

I don't feel very good about the idea of DR just being duplicate servers in a different zone. You don't know what kind of problem could ripple out and affect more zones, or all of them. A completely different host/cloud/colo provider seems like a safer bet.

When you come down to it earth is a single point of failure, its about how big of a disaster you expect to occur and how big of one you expect to recover from.

Yes but it's a lot easier and cheaper to protect against e.g. a major bug surfacing in amazon's cloud management api vs protecting against an asteroid hitting earth.

All the zones are geographically apart so its equivalent to putting server in different cloud/colo and probability that all of them gets "blown" away simultaneously is very less.

And yet all it took was a single AWS availability zone going away for a short while for them to have a major outage.

Major issues in Netflix case per their last blog post due to bugs in their environment not properly failing away from dead ELB's. Also the issues were related to API backups due to everyone rushing to launch new instances in a new AZ, but existing services in other AZ's continued to work fine.

My reading of the Netflix announcement was that it wasn't just a bug, but that they made the conscious decision to include manual intervention in the process (of releasing dead instances) but grossly underestimated the time required to do this across an entire zone.

DR is all about cost, though. It's likely to cost double (or more) to have two (essentially) identical systems in two different cloud providers -- whereas duplicating the same system in two different zones of one provider is nearly free.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.