Skip to content

Comment on Final Root Cause Analysis of Nov 18 Azure Service Interruption

Comments

Nice writeup. I hope that the engineer in question didn't get fired or anything. One of the challenges in SRE/Ops type organizations is to be responsible, take ownership, put in the extra time to fix things you break, but keep the nerve to push out changes. Once an ops team loses its willingness to push large changes, the infrastructure calcifies and you have a much bigger problem on your hand.

I agree, and am reminded of the quote from Thomas Watson (IBM) on this:

“Recently, I was asked if I was going to fire an employee who made a mistake that cost the company $600,000. No, I replied, I just spent $600,000 training him. Why would I want somebody to hire his experience?”

Almost certainly apocryphal. There are multiple versions of the quote floating around. Also: Who was Watson talking to? How did this story get from him to us?

I heard the exact same story in a Wall St. trading firm I worked for - except the story was about a trade gone wrong (trades X shares instead of $X worth of shares, causing the Dow to slide and all sorts of havok).

It just sounds like an anecdote someone would use during a talk.

There's a pretty credible one about AOL's board not firing Steve Case after spending $5 million on him. (For those who have forgotten, Steve didn't actually found AOL; he was brought in from the outside.)

My previous boss used to say that there are two types of engineers - ones that had caused a major outage and ones that never did - and he preferred working with the 1st group as they were naturally more careful.

That's an incomplete analysis.

A more complete one would look at what each engineer had contributed. If the one with the foul-up has and does contribute significant value, and doesn't repeat the same mistakes, it's a good call.

If the careful worker both avoids errors (and costly mistakes) and exceeds other engineers in contributing value, there's a strong argument for keeping them.

There are people who simply foul things up. And there are those who avoid mistakes by simply never taking risks. You almost certainly want to discard the first. The second's value depends on the value your organization gains from innovation.

I would not call it an "analysis" - an moderately insightful joke perhaps, or a heuristic.

At the very beginning of my career, I once made a very stupid and very costly mistake because of tight schedules and lack of proper QA process.

The first thing my boss said the next morning before even explaining what happened was that he was the one responsible and I was not to blame. Oh and he gave me a raise.

Your case is among the 1% luck cases. The other 99% cases are NOT same as yours. They are fired.

I hope it becomes a fireable offence. Any deviation from procedure has to be signed off at the very least.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.