A flat load test and a real daily traffic shape carrying exactly the same total volume, where the flat one stays below capacity and the real one exceeds it at the midday peak

Realistic Traffic Modeling

Back to Performance Tuning and Capacity Planning · Breaking Point Identification · Controlled Environment · Pre-Launch Validation · Service Offerings

Load patterns based on actual usage shapes — bursty, sustained, or spiky — not just a flat ramp-up that misses how real traffic arrives. The illustration above is the whole argument: two days carrying exactly the same number of requests, one of which is a comfortable pass and the other an outage at lunchtime.

1. Capacity Is Consumed by the Peak

Averages are a property of the report; the system experiences the busiest minute. A service handling two million requests a day sounds like 23 per second, and if a third of them arrive in two hours the real figure is several times that.

Measure your peak-to-average ratio once and it will keep being useful. For a consumer service it is commonly three to five; for anything driven by notifications it can be far higher, because everyone is told at the same moment.

2. What a Flat Test Never Exercises

3. The Mix Matters as Much as the Shape

A test that hammers one endpoint measures that endpoint. Production is a mixture, and the expensive requests are usually a small fraction of it — which is precisely why they get left out of the model.

4. Think Time, and the Mistake It Causes

Real users pause between actions. A test without think time has each virtual user hammering continuously, so a hundred virtual users generate the load of several thousand real ones — and the result is read as “we fail at a hundred users”, which is alarming and meaningless.

The honest unit is requests per second, not virtual users. Where user counts are wanted, derive them from the measured session behaviour rather than assuming one request per user per second.

5. The Shape Is Already Recorded

None of this needs to be invented. A week of access logs or a month of request-rate metrics contains the daily curve, the weekly one, the peak-to-average ratio, the endpoint mix and the arrival pattern of your worst burst.

How We Approach It

  1. Extract the real shape from logs or metrics: daily curve, peak-to-average ratio, and the sharpest burst on record.
  2. Extract the real mix, weighted by frequency and including the expensive minority, the writes and the background traffic.
  3. Parameterise the data so virtual users do not all request the same rows.
  4. Model think time from observed sessions, and report in requests per second.
  5. Run the typical peak and the known extreme, and watch recovery after the burst as well as behaviour during it.
  6. Compare the test's own metrics against production's — cache hit rate, error rate, request mix — to confirm the model resembles reality.

What You Get

The question that exposes most load tests: during the test, what was the cache hit rate? If it was far above production's, the database was never really tested.