Illustration of resource traces with one climbing toward a threshold that alerts before the resource is exhausted

Resource Monitoring

Back to Server Setup · Service and Uptime Checks · Capacity Trending · Monitoring Solutions

CPU, memory, disk usage and network throughput tracked continuously, with alerts before thresholds turn into outages. This is the layer that watches the machine. Watching whether the application works is a separate job, covered under service and uptime checks — and a server can be perfectly healthy on every metric here while the service is down.

1. What to Collect, and What It Actually Tells You

2. Set Thresholds That Leave Time to Act

An alert at 95% disk is not an alert, it is a notification of an imminent outage. The threshold should be wherever leaves enough time to do something, which depends on how fast the resource moves and how long the fix takes.

3. Alerts Nobody Reads Are Worse Than None

The most common failure of monitoring is not absence but volume. Once alerts are routinely ignored, the one that mattered is ignored too — and everyone believes they are monitored.

4. Keep Enough History to Answer Questions

Metrics kept for a week answer “what is happening now”. Metrics kept for a year answer the more valuable questions: is this slower than last month, what did the last incident look like, when will we run out. Keep full resolution for the recent period and downsample older data — a year of hourly averages costs very little and is what makes trending possible at all.

5. Put It Somewhere People Look

A dashboard per audience, not one dashboard with everything on it. One overview showing whether things are healthy, and a detailed view per service for when they are not. The test is whether someone unfamiliar can tell, in ten seconds, if anything is wrong — if not, it is a data display rather than a dashboard.

What You Get