
Playbook3 sections · 2 min readBy Dezső Mező · Published October 3, 2026.mdMonitoringOpsUptime
Most monitoring setups fail socially before they fail technically: the metrics exist, the dashboards are pretty, and the outage is still discovered by the customer because the alert went to a channel nobody reads at 9pm. This is the minimal setup that works — three parts, all boring.
01Probe from outside, for the real answer
An external service fetches the real pages and checks the real content — not 'port open' but 'the checkout returns 200 and contains the word pay'. Internal health checks catch dead processes; external probes catch dead certificates, dead DNS and dead upstreams, which is most of what actually dies.
02One alert path that someone owns
Alerts routed to one channel a named person reads — a phone notification, not a shared inbox — with quiet-hours logic that pages anyway for the endpoints that matter. An alert everyone can ignore is an alert everyone does.
03The runbook is written before it is needed
For each alert, a page that says what to check first and what 'fixed' looks like — because the person reading it at 2am is past-you on no sleep. The runbook that lives in someone's head is a single point of failure on a schedule.
What to take away
- External probes check content, not just ports.
- One alert path, one named owner, real notifications.
- Write the runbook at build time, not at 2am.
- Alert fatigue is a bug — every alert must be actionable.
More from the lab
Browse all entries
A backup you never restore is a story
What custom software actually costs
Custom software or off the shelf — the honest test
The first ninety days of a retainer, as it actually goes
Want this looked at on your own system?Start a conversation
