← work

Cutting alert noise by more than 99%

itsm · monitoring · process

The problem

An operations queue at Xbox carried thousands of alerts a day. At that volume an alert queue stops being a signal and becomes background weather. Operators learn to skim, real failures hide inside the noise, and the on-call rotation burns people out on pages that never needed a human.

The method

I ran the same playbook I have used my whole career, one failure mode at a time.

  • Take each alert class and ask the only question that matters: if this fires at 3am, should a human wake up? If the answer is no, the alert gets tuned, rerouted, or retired.
  • Define severity thresholds in writing, in a communication-engagement-notification document, so "how bad is this" stops being a judgment call made differently by every shift.
  • Publish a weekly signal-to-noise report across every operations team, naming each team's noisiest alerts and their trendlines. I ran that report for six months. Nobody wants to top that chart two weeks running, and that is the point.
  • Feed the trendline data back into the ticketing system to enable auto-resolution for the classes of alert that consistently healed themselves.

The outcome

Daily alert volume fell by more than 99%, and every alert that remained was actionable. Average time-to-mitigate came down with it, because operators were reading real signals instead of triaging weather. The reporting habit outlived the project: teams kept watching their own noise because the chart made it visible.

The number is the headline, but the method is the transferable part. Any operation drowning in alerts can run this sequence. It costs discipline, not budget.

If this kind of work is what your team needs, I'm easy to reach.