Load & Capacity Failure Shapes
When a system can’t keep up, the outage usually isn’t “unique.” It tends to follow a small number of repeating shapes. Let's explore some of these shapes to understand why they happen, despite engineers seeing them coming.
1.1 Load Amplification / Retry Storms
This failure shape is what happens when a small failure creates more, new load in an attempt to recover from the failure. A tiny fraction of requests fail, clients retry, those retries become additional traffic, and suddenly the system is overwhelmed by the work it created for itself.
Shape
- A small failure causes retries
- Retries multiply load
- The system collapses under self-inflicted traffic
Typical symptoms
- Some errors recover on retry… until none do
- QPS spikes during outages
- Latency tails explode - the slowest requests get much slower
Bad instincts this triggers
- “If these requests retry again, they might succeed
- “Increase timeouts”
- “Scale everything blindly”
The core idea: retries are load, not healing. If you don’t bound retries with budgets/backoff/circuit breaking, you turn a 2% failure into an incident where your system is now failing because it keeps attempting the same work again and again.
Let’s see what this looks like in a real system.
interactive lab
The system starts returning 500s for ~2% of requests. What do you do?
A small downstream hiccup appears. Nothing is fully down (yet). Your next move determines whether this stays small or turns into self-inflicted load.
Incoming load
55%
Healthy capacity
78%
Error rate
2%
Queue depth
18%
Latency p50
90 ms
Latency p99
260 ms
choose an instinct
1.2 Backpressure Collapse
Backpressure collapse happens when a downstream component slows down but the upstream continues sending work as if nothing changed. The system loses its ability to say “stop” — queues grow, threads block, memory spikes, and the failure becomes a long, slow suffocation.
Shape
- A slow component stops signaling pressure
- Upstream continues sending work
- Queues fill, threads block, memory spikes
Typical symptoms
- Partial hangs (some requests never complete)
- Timeouts without obvious errors
- Long recovery tails even after the fix
Bad instincts this triggers
- “Queue it for later”
- “Add workers”
- “Let it wait”
The core idea: queues convert time problems into space problems. If the system can’t reject, shed, or slow intake, it stores pressure in memory/queue depth — and then fails when it runs out of space.
Let’s see what this looks like in a real system.
interactive lab
A downstream service slows down ~3×. Requests start piling up. What do you do?
Nothing is “down.” But throughput dropped. If upstream keeps pushing at the same rate, pressure has to go somewhere.
Incoming load
60%
Healthy capacity
70%
Error rate
1%
Queue depth
35%
Latency p50
110 ms
Latency p99
420 ms
choose an instinct
1.3 Resource Exhaustion
Resource exhaustion is when a finite resource hits a hard limit: connections, threads, file handles, memory, CPU credits, etc. The system often looks “kind of alive” while it’s actually stuck because all useful capacity is tied up.
Shape
- A finite resource is depleted (connections, threads, file handles)
- Recovery requires draining or restart
- Capacity returns slower than load
Typical symptoms
- Flapping availability
- Hard limits hit
- Partial recovery that regresses
Bad instincts this triggers
- “Increase limits”
- “Add capacity during saturation”
- “Evict idle limits”
The core idea: if you restore capacity without controlling ingress, you often get a fast relapse. Restarts or scaling can “clear” the resource briefly — but the load arrives at the same time, and you hit the limit again.
Let’s see what this looks like in a real system.
interactive lab
DB connection pool is exhausted. Some requests hang. What do you do?
The system isn’t fully dead—it’s stuck. Work is in-flight, but progress is limited because a hard resource cap has been hit.
Incoming load
65%
Healthy capacity
55%
Error rate
3%
Queue depth
40%
Latency p50
140 ms
Latency p99
900 ms
choose an instinct
Match the Symptom to the Shape
A quick diagnostic: given a symptom pattern, which failure shape does it most likely match?
diagnostic mini-quiz
Symptom pattern: QPS spikes during the outage; p99 explodes; errors sometimes recover on retry.