# Monitoring that pages someone, not the dashboard

> A dashboard nobody watches is a screensaver. The setup that actually catches the outage: external probes, one alert path, a written runbook.

Most monitoring setups fail socially before they fail technically: the metrics exist, the dashboards are pretty, and the outage is still discovered by the customer because the alert went to a channel nobody reads at 9pm. This is the minimal setup that works — three parts, all boring.

## Probe from outside, for the real answer

An external service fetches the real pages and checks the real content — not 'port open' but 'the checkout returns 200 and contains the word pay'. Internal health checks catch dead processes; external probes catch dead certificates, dead DNS and dead upstreams, which is most of what actually dies.

## One alert path that someone owns

Alerts routed to one channel a named person reads — a phone notification, not a shared inbox — with quiet-hours logic that pages anyway for the endpoints that matter. An alert everyone can ignore is an alert everyone does.

## The runbook is written before it is needed

For each alert, a page that says what to check first and what 'fixed' looks like — because the person reading it at 2am is past-you on no sleep. The runbook that lives in someone's head is a single point of failure on a schedule.

## What to take away

- External probes check content, not just ports.
- One alert path, one named owner, real notifications.
- Write the runbook at build time, not at 2am.
- Alert fatigue is a bug — every alert must be actionable.

## Tags

Monitoring, Ops, Uptime

## We build this for clients

https://dfieldsolutions.com/en/services/full-stack

## More from the lab

- https://dfieldsolutions.com/en/lab/backup-restore-drill.md — A backup you never restore is a story
- https://dfieldsolutions.com/en/lab/custom-software-cost.md — What custom software actually costs
- https://dfieldsolutions.com/en/lab/custom-vs-off-the-shelf.md — Custom software or off the shelf — the honest test
- https://dfieldsolutions.com/en/lab/first-90-days-retainer.md — The first ninety days of a retainer, as it actually goes

---

Source: https://dfieldsolutions.com/en/lab/monitoring-runbook
DField Solutions — Dunakeszi, Hungary — dezso@dfieldsolutions.com
Booking: see https://dfieldsolutions.com/en/contact
