Spend against headroom, rising gently at first and steeply beyond about sixty per cent, so the last slice of headroom costs several times the first

Cost-Aware Scaling

Back to Performance Tuning and Capacity Planning · Scheduled Cadence · Forecast vs. Actual · Feeds Back Into Tuning · Service Offerings

Balancing headroom against spend, since “just over-provision everything” is not a real capacity strategy. It is a bill that keeps growing while the risk it removes keeps getting smaller — and the growth is not linear.

1. The Shape of the Curve

To carry a load at a given headroom you need capacity equal to the load divided by the fraction you are willing to use. That one over the remaining margin is what produces the curve above.

Below roughly half, extra headroom is cheap and worth buying without much analysis. Past that the curve turns, and each further slice costs more than everything before it. That bend is where the decision actually lives.

2. Headroom Is Not One Number

A single organisation-wide target is almost always wrong in both directions at once. The right target per resource depends on how fast that resource can be added and how badly it fails when exhausted.

3. Price the Buffer Against What It Prevents

Headroom is insurance, and insurance is priced against the loss. Both sides of this can be estimated well enough to decide with.

This turns an argument about caution into an arithmetic comparison, and it occasionally shows that the current buffer is far larger than anything the loss would justify.

4. Cheaper Headroom Than Buying More

Before paying the price on the curve, check whether the same margin is available at a lower one. The alternatives are frequently cheaper by an order of magnitude:

5. Scaling Down Is Part of It

Capacity reviews that only ever add are a ratchet. Load genuinely falls — a client leaves, a feature is retired, an optimisation lands — and the capacity provisioned for it usually stays.

6. Commitment and Elasticity

On metered infrastructure the unit price depends on how it is bought, and the two choices interact with the curve above.

How We Approach It

  1. Set a headroom target per resource, derived from its lead time and its failure behaviour, not one number for everything.
  2. Put a price on the current buffer and on an hour of the outage it prevents.
  3. Replay historical peaks against the proposed thresholds, so frequency is measured rather than assumed.
  4. Exhaust the cheaper headroom first — tuning, caching, shedding, off-peak scheduling, shorter lead times.
  5. Review downwards as well as upwards, in steps, with measurement.
  6. Commit the floor and rent the peak, after verifying the elasticity responds fast enough to matter.

What You Get

The question worth asking: what is the current headroom costing per year, and when was that number last compared with what it prevents?