Home / Problems solved / The outage nobody saw coming
Problem · The outage nobody saw coming

It failed at 2 a.m. on Tuesday. You found out Thursday. From accounting.

The nightly ERP processing quietly exited with an error and every downstream number was wrong for two days. The disk filled and the plant floor found out at the 6 a.m. start. The certificate expired on a Saturday. The backup had been running for three years and nobody had ever restored one. None of these were surprises to the systems. They were only surprises to the people.

What's actually going on

Every outage announces itself first. Nobody was listening.

In every post-mortem we've been part of, the evidence was there beforehand — in a log nobody read, a metric nobody graphed, a job nobody checked. The failure wasn't sudden. The discovery was.

SILENCE

Jobs that fail without telling anyone

The scheduled task exits 1 and nothing routes it to a person. Duration drifts from 40 minutes to 4 hours and nobody graphs it. The failure is logged — to a file nobody opens.

TREND

Failures that were visible a week out

Disk growth rate, queue depth, certificate expiry, backup age. Every one of these is a straight line to a date. Alerting on the failure is too late; alerting on the trend is a calendar entry.

UNTESTED

Recovery that has never been rehearsed

A backup that has never been restored is a hope, not a plan. The first restore should happen on a schedule into a sandbox with a checksum — not on the worst day of the year.

How it gets fixed

Watch the trend. Alert with the evidence. Test the restore.

Install, document, hand over. The goal is that your team sees the problem first — not that you depend on us.

01

Inventory what the business depends on

Every server, service, job, link and integration — and what happens if each one stops. Windows, Linux, Azure, Google, the plant floor. Usually the first time it's been written down.

02

Install the watch

NightOps — a central controller with an agent on every server — for health, night and month-end batch runs, outcome monitors and paging; plus the metrics and logging stack that fits the estate — Azure Monitor / App Insights or Google Cloud Operations for cloud, Prometheus + Grafana or Zabbix / PRTG for servers and network, Loki or Elastic for logs, Sentry for app errors. Tuned so people keep reading the alerts.

03

Alert with the evidence

Every alert carries the log line, the metric and the runbook link. Escalates when ignored. The person paged at 2 a.m. can fix it, not just know about it.

04

Test the restore. Hand over.

Scheduled restores to a sandbox with a checksum. Runbooks written for the on-call person, not for us. Your IT owns it — or a light retainer keeps us on the list.

From the ledger

Built into everything, so the next outage is caught on Monday.

Every system we build ships with this layer in place: the forge shop's quality portal, the Fourth Shift reporting at an agricultural equipment plant, the on-prem Clairvient deployments where the customer's IT policy keeps everything inside their network. Nightly processing is watched for exit codes and duration drift, feeds that go quiet raise alerts, and restores are tested — because a system we built failing silently would be our failure, not the client's.

The same discipline applies to estates we didn't build. The methodology is the same one we've used since a steel mill floor at nineteen, where the furnace didn't wait for anyone to notice: watch what predicts failure, alert with the evidence, rehearse the recovery, and write it down for whoever is on call next.

More entries in the case ledger →

Inventoriedevery dependency the business has, written down — with what happens when it stops
Watchedservers, jobs, network, apps, edges and backups on one screen your IT and your plant manager both read
Actionablealerts with the log line and the runbook attached; the fix is a lookup, not a hunt
Rehearsedrestores tested on a schedule — the only backup test that counts
Let's talk

What went down last, and how did you find out?

That one question tells us most of what we need. Then what runs where — Windows, Linux, Azure, Google, the server under the desk — and whether anyone has ever actually restored a backup.

You'll get a straight answer and a tight scope. If we can't help, we'll tell you that too — and usually who can.

No newsletter. No drip sequence. Just a reply from us.

Got it — thanks.

We read every one of these ourselves. Expect a reply within one business day. If it's urgent, email us directly at dan.mindlin@mindlinconsulting.com.