Circuit Breakers & Bulkheads
How one outage becomes four
Service D slows to 30 seconds per request. Service C calls it and waits, holding a thread and a connection for 30 seconds each time. Under load, C's thread pool fills entirely with requests parked on D. C is now unresponsive — not because anything is wrong with C, but because it is politely waiting.
B calls C and does the same. A calls B. In under a minute, four services are down because one was slow, and only the last one has an actual problem.
Note what the cascade is really made of: waiting. A dependency that fails instantly is far less dangerous than one that hangs, because a fast failure returns your resources immediately. Slow is worse than down.
D slow (30s)
▲
C ───┘ threads block waiting
▲
B ──────┘ threads block waiting
▲
A ─────────┘ threads block waiting
one slow dependency, four dead servicesEvery network call needs a timeout shorter than your caller's patience. An unbounded call is a promise to hold resources for as long as someone else's bad day lasts.
3 components2 connections0:00
Recording…