Skip to content

Comment on Sourcehut network outage post-mortem

Comments

I admire Drew's work and SourceHut, but I wonder if choosing colocation instead of a cloud provider, and not adopting modern cloud tooling hurt the recovery time.

The article mentions how setting up SourceHut from scratch is a complex undertaking with hundreds of small tasks. Some of those are understandable, especially given the amount of data they surely handle, but setting up a complete environment and restoring from a backup should be a simple and mostly automated procedure, not a gargantuan undertaking that needs to involve the entire team several days. There are always difficulties when production is down, and you're trying to restore a full system while undergoing a DDoS attack, I get that, but the reason we have modern cloud providers and tooling is to make creating new environments as painless as possible. It seems foolish not to take advantage of that.

I'm not a fan of Kubernetes either, but it's good that they're experimenting with it. Hopefully it leads to quicker deployments if this happens again.

Cloud isn't a dependency to accomplish this. I can stand up new instances of our primary infrastructure within minutes on any VM or server running linux and our agent.

This sounds like a case of treating your infra as pets and getting stuck when it suddenly needs to be replaced.

That’s not really a surprise considering how they started? To some extend I think it’s unavoidable if you use colocation.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.