Skip to content

Comment on Sourcehut network outage post-mortem

Comments

As unfortunate as these events were, we welcome opportunities to stress-test our emergency procedures;

This right here is invaluable and something you only get from experience. Planning and theory only get you so far.

I extend this thinking to deploying large infrastructure changes you've never done before - you can only plan so much before pulling the trigger and just doing it and seeing what happens.

Wikimedia's operations team go through a full-datacenter failover regularly. That is, once every 6 months or so. It takes several hours of intense all-hands-on-deck operations. They do this repeatedly in order to be sure that all of the procedures are practiced and documentation is well maintained.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.