Parts plotted by how often they fail against how long a replacement takes, with a line at the downtime that can be tolerated; those above it must be on the shelf rather than on a contract

Spare Parts Management

Back to Data Center Management · Warranty Tracking · Scheduled Maintenance Calendar · Escalation Contacts · Service Offerings

This is critical spares on hand for the failures that can't wait on next-day shipping. It exists because a support contract is a promise about time, and some failures have a tolerance shorter than any time the vendor will promise.

1. What to Stock Is an Arithmetic Question

Not "which parts are important" — everything in the rack is important. The question is where the gap is between how long you can be down and how long a replacement takes to arrive under your current contract. Anything on the wrong side of that line goes on the shelf, as in the illustration above.

Which means this decision cannot be made without warranty tracking first: a four-hour onsite contract and a next-business-day contract imply completely different shelves.

2. The Part Everyone Forgets Is the Optic

Disks and power supplies get stocked because they fail visibly and often. The parts that cause the long outages are the ones nobody thought of:

3. A Spare on a Shelf Degrades

Stock is not a one-time purchase. It ages in ways that make it fail at the only moment it is ever used.

Test the spare. Power it, flash it, confirm it is recognised, and put it back. An untested spare is a hypothesis, and the moment you discover it is the wrong revision is the moment you have no replacement at all.

4. Where They Live and Who Knows

5. Failed Hardware Leaves With Your Data

A disk returned under warranty carries whatever was on it. So does a controller with cache, and so does a chassis with a management module holding credentials.

6. Spares Versus Contract Is an Economic Comparison

For commodity hardware the shelf frequently wins. A handful of disks and a power supply can cost less than a year of enhanced cover and respond in minutes rather than hours. For complex or proprietary equipment the contract usually wins, because the part is expensive, rarely needed, and the vendor holds regional stock you cannot match.

The sensible position is usually a mix, decided per part class rather than per vendor, and it is also the answer to some of the renewal decisions in warranty tracking.

How We Approach It

  1. Establish the tolerance for each class of equipment — how long it can actually be down.
  2. Compare that to real replacement time under the contracts you actually hold, not the ones you remember holding.
  3. Produce the stock list from the gap, deliberately including the overlooked categories in section 2.
  4. Set up rotation and testing, so firmware and compatibility stay current.
  5. Register spares as assets, with locations, and make consumption trigger replenishment.
  6. Settle the data destruction position for failed media and returned equipment.

What You Get

The test is uncomfortable and quick: pick the part you have never replaced, and ask how long it would take to get one tonight.