Threshold Alerting
Back to Data Center Management · Sensor Coverage · Power Draw Per Circuit · Historical Trending · Service Offerings
Threshold alerting means notifications before conditions drift out of safe range, not after equipment starts failing. Every part of that sentence is a design constraint, and the one people get wrong is the timing: an alert set at the equipment's limit is not an early warning. It is a notification that you have already run out of room.
1. Set the Threshold Backwards From the Limit
The correct order is to start from the number you must not reach, subtract the time you need to do something about it, and put the alert there.
- The hard limit. The equipment's rated inlet maximum, the 80 per cent continuous figure for a circuit — see power draw per circuit — whatever the physical constraint actually is.
- The ride-through. How long from "cooling has failed" to "the limit is reached". In a dense rack with containment this can be a very small number: a few minutes, not an hour. Measure it if you can, by watching what happens during a planned cooling maintenance.
- The response time. How long to notice, decide and act — which at 3am on a Sunday includes someone waking up and possibly driving somewhere.
- The warning threshold goes far enough below the limit that steps two and three both fit inside the gap. That gap is the window shown in the illustration above.
If the arithmetic does not work — if the ride-through is shorter than the response time — then no threshold will save you, and the honest conclusion is that the problem needs fixing somewhere other than alerting: more cooling redundancy, lower density, or someone closer.
2. Two Levels, With Different Meanings
| Level | Means | Goes to |
|---|---|---|
| Warning | Something has changed and there is still time. Investigate. | A queue someone reads during the working day |
| Critical | Act now, or equipment is at risk. | Whoever is on call, by a method that wakes them |
Keeping these distinct is what keeps the critical level credible. The moment warnings start paging people at night, the pages stop being read — and the first one that mattered is lost with the rest. See escalation path for where each level should land.
3. Rate of Change Beats Absolute Value
The single most useful environmental alert is not a temperature. It is a slope.
If a rack's inlet rises by one degree per minute, something has failed — a cooling unit, a fan, airflow. You know that within sixty seconds, while the absolute temperature is still perfectly normal and no threshold has been approached. Waiting for the absolute value throws away the entire head start.
- Temperature rising faster than X per minute catches cooling failure almost immediately.
- A step change in power draw means something started or stopped that you did not expect.
- Delta-T collapsing means air has stopped moving through the equipment, which precedes the temperature problem it will cause.
Rate alerts and absolute alerts answer different questions and you want both: the slope tells you something broke, the level tells you how much time is left.
4. Alert on Silence
A sensor that stops reporting produces no alerts, and a dashboard with no alerts looks like a healthy room. This is the most dangerous quiet failure in the whole chapter, and it is trivially avoidable: treat stale data as a condition in its own right.
- No reading from a sensor for longer than N intervals raises a warning.
- The alerting path itself is monitored — a heartbeat that fires if the pipeline stops.
- Sensors that fail loudly are preferred to sensors that repeat their last good value, because a frozen plausible number defeats every other check on this page.
5. Stopping the Noise Without Losing the Signal
- Hysteresis. Clear the alert at a lower value than the one that raised it, otherwise a reading sitting on the line flaps repeatedly.
- Sustained for a duration. Require the condition to hold for a few minutes. A single sample over the line is usually a sensor artefact; three minutes over it is a room.
- Group related alerts. A cooling failure trips every rack in the row; that is one incident and should arrive as one page, with the detail attached.
- Do not alert on what nobody will act on. Every alert that reliably produces no action trains people to ignore the channel it arrives on.
6. An Untested Alert Is a Hypothesis
Configuring an alert and seeing it in the list is not evidence that it works. The path has several links — sensor, collector, rule, notification provider, the recipient's phone and its do-not-disturb settings — and any one of them can be broken silently for months.
- Trigger each critical alert deliberately and confirm it arrives, including out of hours.
- Repeat after changes to monitoring, on-call rotas or notification providers.
- Record who received it and how long it took. That figure is the response time in section 1, and it is usually worse than assumed.
- Fold the findings into post-incident follow-up, which is where thresholds get corrected by experience.
How We Approach It
- Collect the hard limits for every monitored thing — inlet maxima, circuit ratings, humidity envelope.
- Establish ride-through, measured where possible rather than assumed.
- Measure the real response time, including out of hours.
- Derive the warning levels from those three numbers, and set critical at the point where action can no longer wait.
- Add rate-of-change and staleness alerts, which catch what levels cannot.
- Apply hysteresis, durations and grouping, then route warning and critical to genuinely different destinations.
- Test every critical path end to end and record the result.
What You Get
- Warning and critical thresholds derived from limits and response time, with the reasoning written down rather than inherited from a default.
- Rate-of-change alerts that detect a cooling failure in its first minute.
- Staleness alerting, so a silent sensor is not mistaken for a quiet room.
- Hysteresis, durations and grouping tuned so that a single event produces a single page.
- Every critical alert tested end to end out of hours, with the measured delivery time recorded.
A good threshold is one that has woken somebody up while there was still something useful for them to do. Anything that fires later is a record of an outage, not an alert.