Skip to main content

# lab note

Alert hygiene: dashboards that page you only when it matters

SLO-based alerting cut our noisy pages by 80%. The config and the reasoning.

1 min read1 Script = 1 Problem Gone

Sean · wrote this at the bench

observabilitygrafanaslo

Every alert that pages you and isn’t actionable trains you to ignore the next one. We rebuilt alerting around error budgets and burn rate.

The problem

A small, annoying task that was eating time and attention every week. We measured it before touching anything.

The approach

One focused change, reproducible, with a guardrail and an obvious way to turn it off. No magic, no vibes.

$ ./run --dry-run
[ ok ] plan looks sane
$ ./run --apply

The receipts

Before/after benchmarks, the failure modes we hit, and the tradeoffs we accepted. The repo has the full runbook.

$ subscribe --email

Automation patterns and lab updates. No hype, no spam.