Scheduled Maintenance
Back to Server Setup · Patch Management · Backup and Recovery · Logrotate Builder
OS and security patching, log rotation, backup verification, and certificate renewals on a set cadence, not “whenever someone remembers.” None of these is difficult. All of them fail the same way: they are important and never urgent, so they lose to whatever is urgent, until one of them becomes an outage.
1. Patching, on a Rhythm
- Security updates automatically, as set up under security settings. This covers most of the exposure with no human in the loop.
- A monthly window for the rest — non-security updates, firmware, anything needing a reboot. Predictable and announced, so the business can plan around it.
- Non-production first, always. The same patches, a few days earlier, somewhere it does not matter.
- An out-of-cycle path for a critical vulnerability under active exploitation, decided in advance rather than improvised. The detail belongs to patch management.
- Reboots are part of it. A kernel patched and never booted is not applied, and a server with six months of uptime has six months of unapplied kernel fixes. Regular reboots also prove the machine still comes back, which is worth knowing before you find out during an incident.
2. Log Rotation, Before the Disk Fills
An unrotated log is a scheduled outage with a long fuse. Rotation by size as well as by age — a daily rotation does not help if something logs twenty gigabytes in an hour — compression of old files, and a retention that matches what aggregation already keeps centrally, since the copy on the host is only a buffer. Our logrotate builder generates the configuration.
The detail that catches people: a rotated file the application keeps writing into. Without the right signal or a copy-and-truncate approach, the process holds the deleted file open and the space is never returned — the disk fills while every file on it looks small.
3. Backup Verification, Which Means Restoring
A backup job reporting success proves a job ran. It does not prove the data is complete, that the archive is readable, or that anyone knows the procedure. The only evidence is a restore.
- An automated restore test, regularly, to a scratch environment, with an assertion on the restored data rather than only on the exit code.
- A manual full restore periodically, timed, so the recovery objective is a measured number rather than an aspiration.
- Someone other than the usual person doing it occasionally, following the written runbook. That is how you find out the runbook is incomplete.
- Check coverage, not just success — a new database added six months ago that nobody added to the backup job is the classic finding.
4. Certificates and Other Expiring Things
Automate renewal, monitor it, and alert on approaching expiry independently of the automation — because the failure mode is the renewal silently not working. Then extend the same thinking: domain registrations, support contracts, cloud commitments, signing keys, API credentials with an expiry, and the licences covered by the licensing check. Anything with a date attached belongs on one list with an owner.
5. Cleanup and Drift
A few jobs that prevent slow decay: removing old container images and build artefacts, which consume disk invisibly; pruning obsolete snapshots; checking filesystem growth against the trend from capacity trending; and re-running configuration management in check mode to see what has been changed by hand since the last run. Systems drift away from their intended state constantly, and the only question is whether you find out deliberately or during an incident.
6. Automate It, Then Watch the Automation
Everything here should run without a person, and every one of them should report. The specific risk of automated maintenance is silent failure: a backup job that has been erroring for three weeks, a renewal that stopped working, a patch run that fails on one host. Each job reports success, and the absence of a report is itself an alert. A calendar of what runs when, with an owner per item, is the deliverable — not the individual scripts.
What You Get
- A maintenance calendar: what runs, how often, in which window, and who owns it.
- Automatic security patching plus a monthly window for the rest, non-production first.
- Log rotation configured by size and age, with the held-open-file case handled.
- Automated restore tests, and periodic timed full restores that produce a real recovery figure.
- One list of everything that expires, with owners and independent alerting.
- Reporting on every job, and alerts when a report does not arrive.