Home /Failure shapes

Load & Capacity Failure Shapes

When a system can’t keep up, the outage usually isn’t “unique.” It tends to follow a small number of repeating shapes. Let's explore some of these shapes to understand why they happen, despite engineers seeing them coming.

1.1 Load Amplification / Retry Storms

This failure shape is what happens when a small failure creates more, new load in an attempt to recover from the failure. A tiny fraction of requests fail, clients retry, those retries become additional traffic, and suddenly the system is overwhelmed by the work it created for itself.

Shape

  • A small failure causes retries
  • Retries multiply load
  • The system collapses under self-inflicted traffic

Typical symptoms

  • Some errors recover on retry… until none do
  • QPS spikes during outages
  • Latency tails explode - the slowest requests get much slower

Bad instincts this triggers

  • “If these requests retry again, they might succeed
  • “Increase timeouts”
  • “Scale everything blindly”

The core idea: retries are load, not healing. If you don’t bound retries with budgets/backoff/circuit breaking, you turn a 2% failure into an incident where your system is now failing because it keeps attempting the same work again and again.

Let’s see what this looks like in a real system.

interactive lab

The system starts returning 500s for ~2% of requests. What do you do?

A small downstream hiccup appears. Nothing is fully down (yet). Your next move determines whether this stays small or turns into self-inflicted load.

Client → API → Worker Pool → DB

Incoming load

55%

Healthy capacity

78%

Error rate

2%

Queue depth

18%

Latency p50

90 ms

Latency p99

260 ms

choose an instinct

1.2 Backpressure Collapse

Backpressure collapse happens when a downstream component slows down but the upstream continues sending work as if nothing changed. The system loses its ability to say “stop” — queues grow, threads block, memory spikes, and the failure becomes a long, slow suffocation.

Shape

  • A slow component stops signaling pressure
  • Upstream continues sending work
  • Queues fill, threads block, memory spikes

Typical symptoms

  • Partial hangs (some requests never complete)
  • Timeouts without obvious errors
  • Long recovery tails even after the fix

Bad instincts this triggers

  • “Queue it for later”
  • “Add workers”
  • “Let it wait”

The core idea: queues convert time problems into space problems. If the system can’t reject, shed, or slow intake, it stores pressure in memory/queue depth — and then fails when it runs out of space.

Let’s see what this looks like in a real system.

interactive lab

A downstream service slows down ~3×. Requests start piling up. What do you do?

Nothing is “down.” But throughput dropped. If upstream keeps pushing at the same rate, pressure has to go somewhere.

Client → API → Queue → Workers → Provider ↓ slower x3

Incoming load

60%

Healthy capacity

70%

Error rate

1%

Queue depth

35%

Latency p50

110 ms

Latency p99

420 ms

choose an instinct

1.3 Resource Exhaustion

Resource exhaustion is when a finite resource hits a hard limit: connections, threads, file handles, memory, CPU credits, etc. The system often looks “kind of alive” while it’s actually stuck because all useful capacity is tied up.

Shape

  • A finite resource is depleted (connections, threads, file handles)
  • Recovery requires draining or restart
  • Capacity returns slower than load

Typical symptoms

  • Flapping availability
  • Hard limits hit
  • Partial recovery that regresses

Bad instincts this triggers

  • “Increase limits”
  • “Add capacity during saturation”
  • “Evict idle limits”

The core idea: if you restore capacity without controlling ingress, you often get a fast relapse. Restarts or scaling can “clear” the resource briefly — but the load arrives at the same time, and you hit the limit again.

Let’s see what this looks like in a real system.

interactive lab

DB connection pool is exhausted. Some requests hang. What do you do?

The system isn’t fully dead—it’s stuck. Work is in-flight, but progress is limited because a hard resource cap has been hit.

API → Worker Pool → DB (threads) (conn pool maxed)

Incoming load

65%

Healthy capacity

55%

Error rate

3%

Queue depth

40%

Latency p50

140 ms

Latency p99

900 ms

choose an instinct

Match the Symptom to the Shape

A quick diagnostic: given a symptom pattern, which failure shape does it most likely match?

diagnostic mini-quiz

Symptom pattern: QPS spikes during the outage; p99 explodes; errors sometimes recover on retry.