A rising temperature trace crossing a warning threshold well before the critical limit, with the interval between the two marked as the window available to act

Threshold Alerting

Back to Data Center Management · Sensor Coverage · Power Draw Per Circuit · Historical Trending · Service Offerings

Threshold alerting means notifications before conditions drift out of safe range, not after equipment starts failing. Every part of that sentence is a design constraint, and the one people get wrong is the timing: an alert set at the equipment's limit is not an early warning. It is a notification that you have already run out of room.

1. Set the Threshold Backwards From the Limit

The correct order is to start from the number you must not reach, subtract the time you need to do something about it, and put the alert there.

  1. The hard limit. The equipment's rated inlet maximum, the 80 per cent continuous figure for a circuit — see power draw per circuit — whatever the physical constraint actually is.
  2. The ride-through. How long from "cooling has failed" to "the limit is reached". In a dense rack with containment this can be a very small number: a few minutes, not an hour. Measure it if you can, by watching what happens during a planned cooling maintenance.
  3. The response time. How long to notice, decide and act — which at 3am on a Sunday includes someone waking up and possibly driving somewhere.
  4. The warning threshold goes far enough below the limit that steps two and three both fit inside the gap. That gap is the window shown in the illustration above.

If the arithmetic does not work — if the ride-through is shorter than the response time — then no threshold will save you, and the honest conclusion is that the problem needs fixing somewhere other than alerting: more cooling redundancy, lower density, or someone closer.

2. Two Levels, With Different Meanings

Level Means Goes to
Warning Something has changed and there is still time. Investigate. A queue someone reads during the working day
Critical Act now, or equipment is at risk. Whoever is on call, by a method that wakes them

Keeping these distinct is what keeps the critical level credible. The moment warnings start paging people at night, the pages stop being read — and the first one that mattered is lost with the rest. See escalation path for where each level should land.

3. Rate of Change Beats Absolute Value

The single most useful environmental alert is not a temperature. It is a slope.

If a rack's inlet rises by one degree per minute, something has failed — a cooling unit, a fan, airflow. You know that within sixty seconds, while the absolute temperature is still perfectly normal and no threshold has been approached. Waiting for the absolute value throws away the entire head start.

Rate alerts and absolute alerts answer different questions and you want both: the slope tells you something broke, the level tells you how much time is left.

4. Alert on Silence

A sensor that stops reporting produces no alerts, and a dashboard with no alerts looks like a healthy room. This is the most dangerous quiet failure in the whole chapter, and it is trivially avoidable: treat stale data as a condition in its own right.

5. Stopping the Noise Without Losing the Signal

6. An Untested Alert Is a Hypothesis

Configuring an alert and seeing it in the list is not evidence that it works. The path has several links — sensor, collector, rule, notification provider, the recipient's phone and its do-not-disturb settings — and any one of them can be broken silently for months.

How We Approach It

  1. Collect the hard limits for every monitored thing — inlet maxima, circuit ratings, humidity envelope.
  2. Establish ride-through, measured where possible rather than assumed.
  3. Measure the real response time, including out of hours.
  4. Derive the warning levels from those three numbers, and set critical at the point where action can no longer wait.
  5. Add rate-of-change and staleness alerts, which catch what levels cannot.
  6. Apply hysteresis, durations and grouping, then route warning and critical to genuinely different destinations.
  7. Test every critical path end to end and record the result.

What You Get

A good threshold is one that has woken somebody up while there was still something useful for them to do. Anything that fires later is a record of an outage, not an alert.