Skip to content

Comment on Final Root Cause Analysis of Nov 18 Azure Service Interruptionparent

Comments

My previous boss used to say that there are two types of engineers - ones that had caused a major outage and ones that never did - and he preferred working with the 1st group as they were naturally more careful.

That's an incomplete analysis.

A more complete one would look at what each engineer had contributed. If the one with the foul-up has and does contribute significant value, and doesn't repeat the same mistakes, it's a good call.

If the careful worker both avoids errors (and costly mistakes) and exceeds other engineers in contributing value, there's a strong argument for keeping them.

There are people who simply foul things up. And there are those who avoid mistakes by simply never taking risks. You almost certainly want to discard the first. The second's value depends on the value your organization gains from innovation.

I would not call it an "analysis" - an moderately insightful joke perhaps, or a heuristic.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.