Illustration of a day of traffic with a sharp peak below a headroom band, beside CPU, memory, disk and network meters with memory marked as the constraint

Workload Requirements

Back to Server Setup · Capacity Planning · Service Offerings

Defining workload requirements means establishing what applications and services need to run, the expected traffic and usage patterns, and the peak-load scenarios — before anyone chooses hardware. It is the first stage of a server setup, and skipping it is why so many servers are simultaneously oversized and too slow: too much of the resource that was never the constraint, too little of the one that was.

The useful question is never “how big a server do we need?”. It is four questions: what work, how much of it, how unevenly it arrives, and how fast it has to be answered. Get those, and the hardware follows almost mechanically. Skip them, and you are buying by intuition.

1. What the Workload Actually Needs

Every workload is constrained by one resource at a time, and adding more of any other changes nothing. The first job is to identify which.

Most real systems are one of these with a secondary constraint waiting behind it. Identifying the order matters, because relieving the first one promotes the second — which is why an upgrade sometimes produces less improvement than expected. Where a system already exists, this is measurement, not estimation: see request profiling.

Beyond the four, we record what the software itself demands: supported operating systems and versions, minimum specifications the vendor will actually support, GPU requirements, architecture (x86-64 or ARM) and whether every dependency has builds for it, and licence constraints — which can make the core count a commercial decision rather than a technical one.

2. The Shape of the Traffic

Averages hide everything that matters. A server sized for mean load will be underwater during every peak that mean was averaged from.

3. Peak-Load Scenarios

The peak that breaks a system is usually not the busy hour. It is a specific, nameable event, and the exercise is to list them explicitly rather than to add a safety margin and hope.

4. Why Headroom Is Arithmetic, Not Caution

“We're only at 85% CPU” sounds like efficient use of a purchase. It is not, and the reason is queueing rather than opinion. As utilisation rises, requests increasingly wait behind other requests, and that wait grows without limit as utilisation approaches 100%.

Chart of response time as a multiple of service time against utilisation: flat below 70 percent, five times slower at 80 percent, ten times slower at 90 percent

In the idealised single-queue model, response time is service time divided by (1 − utilisation). At 50% busy, work takes twice as long as the work itself. At 80%, five times. At 90%, ten times. At 95%, twenty. Real systems deviate in the details — multiple cores, uneven arrivals and variable service times all shift the curve — but none of them removes the bend, and bursty traffic makes it worse rather than better.

The practical consequence: it is peak utilisation that has to sit left of the bend, not average. A server averaging 30% but hitting 92% for twenty minutes every afternoon is a server that is slow every afternoon, and the daily average will never show it.

5. The Requirements That Are Not About Load

6. When There Is Nothing to Measure Yet

For a system that does not exist, honest estimation beats confident guessing:

What You Get

This feeds straight into server selection and, once the system is live, into capacity planning — where the same numbers get compared against reality.