Illustration of a maintenance calendar with recurring tasks marked, beside the four routine jobs and the cost of doing them ad hoc

Scheduled Maintenance

Back to Server Setup · Patch Management · Backup and Recovery · Logrotate Builder

OS and security patching, log rotation, backup verification, and certificate renewals on a set cadence, not “whenever someone remembers.” None of these is difficult. All of them fail the same way: they are important and never urgent, so they lose to whatever is urgent, until one of them becomes an outage.

1. Patching, on a Rhythm

2. Log Rotation, Before the Disk Fills

An unrotated log is a scheduled outage with a long fuse. Rotation by size as well as by age — a daily rotation does not help if something logs twenty gigabytes in an hour — compression of old files, and a retention that matches what aggregation already keeps centrally, since the copy on the host is only a buffer. Our logrotate builder generates the configuration.

The detail that catches people: a rotated file the application keeps writing into. Without the right signal or a copy-and-truncate approach, the process holds the deleted file open and the space is never returned — the disk fills while every file on it looks small.

3. Backup Verification, Which Means Restoring

A backup job reporting success proves a job ran. It does not prove the data is complete, that the archive is readable, or that anyone knows the procedure. The only evidence is a restore.

4. Certificates and Other Expiring Things

Automate renewal, monitor it, and alert on approaching expiry independently of the automation — because the failure mode is the renewal silently not working. Then extend the same thinking: domain registrations, support contracts, cloud commitments, signing keys, API credentials with an expiry, and the licences covered by the licensing check. Anything with a date attached belongs on one list with an owner.

5. Cleanup and Drift

A few jobs that prevent slow decay: removing old container images and build artefacts, which consume disk invisibly; pruning obsolete snapshots; checking filesystem growth against the trend from capacity trending; and re-running configuration management in check mode to see what has been changed by hand since the last run. Systems drift away from their intended state constantly, and the only question is whether you find out deliberately or during an incident.

6. Automate It, Then Watch the Automation

Everything here should run without a person, and every one of them should report. The specific risk of automated maintenance is silent failure: a backup job that has been erroring for three weeks, a renewal that stopped working, a patch run that fails on one host. Each job reports success, and the absence of a report is itself an alert. A calendar of what runs when, with an owner per item, is the deliverable — not the individual scripts.

What You Get