Connections rising past a file descriptor limit of 1024, after which every further connection is refused while nothing on the machine is overloaded

Kernel and Filesystem Parameters

Back to Performance Tuning and Capacity Planning · Connection Pool Sizing · Caching Strategy · Before-and-After Measurement · Service Offerings

File descriptor limits, TCP settings, and I/O schedulers adjusted for actual workload instead of distro defaults sized for a generic server. The defaults are not wrong; they are a compromise for a machine whose job nobody knew. Yours has a job.

This is the parameter-level companion to performance tuning at build time, which covers the method and the wider picture. Here the subject is the specific settings, which default bites which workload, and how to know before it does.

1. The Defaults That Become Walls

What makes these dangerous is that the machine looks healthy while they bite. CPU is low, memory is fine, the disk is idle — and connections are being refused, as in the illustration above. There is no gradual degradation to notice.

Limit What happens when you reach it Who hits it
Open file descriptors Accepts fail, log writes fail, the application reports errors that read as bugs Anything holding many concurrent connections — a proxy, a websocket server, a busy database client
Listen backlog (somaxconn) Connections dropped during bursts, seen by the client as a timeout rather than a refusal Services with spiky arrival patterns
Ephemeral port range Outbound connections fail once the range is exhausted by sockets in TIME_WAIT A service making many short-lived outbound calls to one destination
Connection tracking table Packets silently dropped once full, with a kernel log line nobody is watching Any host with a stateful firewall and high connection rates
Inotify watches File watchers stop working, usually in a build tool, usually blamed on the tool Development machines and CI runners

2. The File Descriptor Limit Is Three Limits

Raising it is the most common tuning task and the most commonly done incompletely, because there are several in series and the smallest wins.

That last one is the trap. Raising the limit in a configuration file, confirming it in a terminal, and restarting the service changes nothing, because the service was never started from that terminal. Check the running process rather than the shell.

3. Match the I/O Scheduler to the Device

Schedulers exist to reorder requests so a mechanical head travels less. On a device with no head, that work is pure overhead.

4. Filesystem Settings Worth Checking

5. TCP Settings, and the Ones to Leave Alone

Network tuning attracts copied configuration more than any other area, and most of it is a decade out of date. Modern kernels ship sensible values, and several once-popular settings are now harmful or removed.

6. A Setting That Does Not Survive a Reboot Is Not a Fix

A value applied at the command line disappears, and it disappears months later, during a restart nobody connected to the symptom. The change belongs in configuration, with three things alongside it: what it was, why it was changed, and what measurement justified it.

Where the machine is built by automation, the parameter belongs there rather than on the machine — see infrastructure as code. The failure mode otherwise is a host that was tuned, rebuilt, and quietly went back to defaults.

How We Approach It

  1. Read the current effective values from the running processes, not from the files that are supposed to set them.
  2. Compare against the workload: peak concurrent connections against the descriptor limit, outbound connection rate against the port range, connection rate against the tracking table.
  3. Check the storage settings against the actual device, measured rather than assumed, which matters most on virtualised disks.
  4. Change what the measurements justify, one at a time, and leave the rest alone.
  5. Persist every change in configuration or in the build, with the reason recorded next to it.
  6. Add alerting on the limits themselves, so the next one announces itself instead of arriving as a refused connection.

What You Get

The symptom to recognise: a machine with plenty of everything, refusing work. That is almost never capacity. It is a number someone chose for a different server.