The nightly ERP processing quietly exited with an error and every downstream number was wrong for two days. The disk filled and the plant floor found out at the 6 a.m. start. The certificate expired on a Saturday. The backup had been running for three years and nobody had ever restored one. None of these were surprises to the systems. They were only surprises to the people.
In every post-mortem we've been part of, the evidence was there beforehand — in a log nobody read, a metric nobody graphed, a job nobody checked. The failure wasn't sudden. The discovery was.
The scheduled task exits 1 and nothing routes it to a person. Duration drifts from 40 minutes to 4 hours and nobody graphs it. The failure is logged — to a file nobody opens.
Disk growth rate, queue depth, certificate expiry, backup age. Every one of these is a straight line to a date. Alerting on the failure is too late; alerting on the trend is a calendar entry.
A backup that has never been restored is a hope, not a plan. The first restore should happen on a schedule into a sandbox with a checksum — not on the worst day of the year.
Install, document, hand over. The goal is that your team sees the problem first — not that you depend on us.
Every server, service, job, link and integration — and what happens if each one stops. Windows, Linux, Azure, Google, the plant floor. Usually the first time it's been written down.
NightOps — a central controller with an agent on every server — for health, night and month-end batch runs, outcome monitors and paging; plus the metrics and logging stack that fits the estate — Azure Monitor / App Insights or Google Cloud Operations for cloud, Prometheus + Grafana or Zabbix / PRTG for servers and network, Loki or Elastic for logs, Sentry for app errors. Tuned so people keep reading the alerts.
Every alert carries the log line, the metric and the runbook link. Escalates when ignored. The person paged at 2 a.m. can fix it, not just know about it.
Scheduled restores to a sandbox with a checksum. Runbooks written for the on-call person, not for us. Your IT owns it — or a light retainer keeps us on the list.
Every system we build ships with this layer in place: the forge shop's quality portal, the Fourth Shift reporting at an agricultural equipment plant, the on-prem Clairvient deployments where the customer's IT policy keeps everything inside their network. Nightly processing is watched for exit codes and duration drift, feeds that go quiet raise alerts, and restores are tested — because a system we built failing silently would be our failure, not the client's.
The same discipline applies to estates we didn't build. The methodology is the same one we've used since a steel mill floor at nineteen, where the furnace didn't wait for anyone to notice: watch what predicts failure, alert with the evidence, rehearse the recovery, and write it down for whoever is on call next.
That one question tells us most of what we need. Then what runs where — Windows, Linux, Azure, Google, the server under the desk — and whether anyone has ever actually restored a backup.
You'll get a straight answer and a tight scope. If we can't help, we'll tell you that too — and usually who can.