Skip to content

Comment on Slack’s Outage on January 4th 2021

Comments

our dashboarding and alerting service became unavailable.

Sounds like the monitoring system needs a monitoring system.

It is quite awkward that the output of "working" and "completely broken" alerting systems have the same visible effect -- no alerts.

For Prometheus users, I wrote alertmanager-status to let a third-party "website up?" monitoring server check your alertmanager: https://github.com/jrockway/alertmanager-status

(I also wrote one of the main Google Fiber monitoring systems back when I was at Google. We spent quite a bit of time on monitoring monitoring, because whenever there was an actual incident people would ask us "is this real, or just the monitoring system being down?" Previous monitoring systems were flaky so people were kind of conditioned to ignore the improved system -- so we had to have a lot of dashboards to show them that there was really an ongoing issue.)

"Who monitors the monitors?"

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.