The check that matters is from outside: a monitor on the same server reports 'all good' right up until the server dies. External probes — HTTP 200 with the right content, not just a ping — catch the real failures: expired certificates, dead app servers, a deploy that returns the maintenance page.
The honest target setting matters more than the tool: four nines (99.99%) is fifty-two minutes of downtime a year; three is eight hours. Monitoring with no on-call answer is a log of when you were down, not protection — the alert has to reach someone who can act.
Related terms
Incident response
The plan and the practice for what happens between noticing something is wrong and being back to normal.
HTTP caching
Headers that tell browsers and CDNs how long a response may be reused — the cheapest performance optimization that exists.
CI/CD
Automatically building and testing every change, and automatically shipping the ones that pass.
The bench this belongs to
Full-stackIf you can describe it, we can build it. React and Next.js on the front, Python or Node behind, Postgres underneath, shipped to somewhere you can afford to run.
