Sensor Coverage
Back to Data Center Management · Power Draw Per Circuit · Threshold Alerting · Historical Trending · Service Offerings
Sensor coverage means temperature and humidity monitored at rack level, not just one reading for the whole room. The distinction is not a refinement. A single room reading is an average, and an average describes a place that does not exist — it is comfortably within range while one rack cooks.
1. Measure at the Intake, Not in the Aisle
Equipment is rated on the temperature of the air entering it. That is the number the manufacturer's warranty refers to and the number that determines whether a machine throttles. So the sensor belongs at the air intake — the front face of the rack — and not wherever there happened to be a convenient mounting point.
Three heights per rack, at the top, middle and bottom of the intake face. The top is almost always the warmest, because hot exhaust recirculates over the top of the rack and because the cold supply has furthest to travel. A single mid-height sensor misses exactly the reading you most need. This is what the illustration above shows: the same rack reads 19 °C at the bottom and 28 °C at the top.
ASHRAE's TC 9.9 thermal guidelines are the usual reference point — a recommended inlet envelope of roughly 18–27 °C, with wider allowable classes. Whatever envelope you adopt, apply it to intake readings, because applying it to a room average means nothing.
2. What Else Is Worth Instrumenting
- Rack exhaust, at least on a sample. The difference between intake and exhaust (delta-T) tells you whether air is actually moving through the equipment. A delta-T that collapses means airflow has stopped somewhere, which is a problem before it is a temperature.
- Cooling supply and return. At the CRAC or in-row unit. A unit whose return temperature is falling is short-cycling — cooling the room's own cold air rather than equipment exhaust — which wastes capacity you have already paid for.
- Humidity, as dew point. Relative humidity varies with temperature and so reads differently at two points in the same room; dew point does not, which makes it the honest control variable. Too dry raises electrostatic discharge risk; too damp risks condensation and corrosion.
- Leak detection wherever water is: under a raised floor, beneath in-row coolers, near pipework and humidifiers. Rope sensors rather than spot sensors, because water goes where it wants.
- Differential pressure if you have aisle containment or an underfloor plenum — it tells you whether the containment is doing anything.
- Door position on containment and cabinets, since a propped-open door quietly undoes the containment and nothing else will report it.
3. Some Coverage You Already Own
Before buying sensors, collect what the estate already reports. Most of it is free and nobody reads it.
- Servers report their own inlet temperature over IPMI or Redfish. That is a sensor at the exact point that matters, already fitted, on every machine. It is the single cheapest improvement in coverage available to most sites.
- Intelligent PDUs commonly take temperature and humidity probes, and sit at the back of the rack already.
- Switches and storage controllers expose inlet and component temperatures over SNMP.
Pulling these into your existing monitoring gives per-rack coverage without a procurement cycle. Dedicated sensors then fill the gaps: aisles, containment, floor voids, and anywhere without powered equipment to ask.
4. The Monitoring Must Not Depend on What It Monitors
This is the failure that turns a warm aisle into a lost rack. If the sensors report through a switch in the rack they are watching, or the collector runs on a server in that room, then a thermal event takes out the monitoring at the same moment it takes out the equipment — and the alert you were relying on never arrives. The graph simply stops, which is easy to read as "nothing happening".
- Sensor gateways on their own power feed, ideally the UPS-backed one.
- Reporting out of the room over a path that does not depend on the room.
- Alerting on missing data as well as on bad data — a sensor that stops reporting must raise something. See threshold alerting.
- Sensors that fail loudly rather than repeating their last good value, which is the worst possible failure mode because everything looks fine.
5. Placement, Calibration and the Boring Parts
- Record where every sensor is, to the rack and the height, as an asset in its own right. A reading without a known location is not usable. This is why sensors belong in the asset register and on the rack elevation.
- Not in the exhaust path of something else, not above a PDU, and not where a cable bundle has grown over it.
- Calibrate or at least cross-check on a schedule. Two sensors a metre apart that disagree by three degrees mean one of them is lying and you do not know which.
- Re-survey after every layout change. Airflow is a property of the room as arranged, and adding a rack changes it.
How We Approach It
- Inventory what already reports — server inlets over IPMI or Redfish, PDU probes, switch sensors. This usually covers more than expected.
- Survey the room for hot spots and recirculation, which identifies where the gaps actually are rather than spreading sensors evenly.
- Place intake sensors at three heights on every rack that matters, plus exhaust sampling, cooling supply and return, and leak rope wherever there is water.
- Make it independent — separate power, a reporting path that survives the room, and alerting on silence.
- Record every sensor's location in the asset register and on the elevations.
- Set the envelope against intake readings, and hand the numbers to threshold alerting and trending.
What You Get
- Intake temperature at three heights per rack, with existing on-board sensors used before anything is bought.
- Humidity tracked as dew point, plus leak detection wherever water is present.
- Delta-T and cooling supply and return readings, so airflow problems are visible as airflow problems.
- A monitoring path that does not fail with the room, and alerts when a sensor goes quiet.
- Every sensor located in the asset register and on the rack elevation, with a cross-check schedule.
The test is straightforward: if a cooling unit failed at 2am on a Sunday, would you know which rack was in trouble, or only that the room average had moved a little?