Resilience patterns
Stop one failing service from taking down the rest with timeouts, retries with backoff, circuit breakers, bulkheads and fallbacks.
- Explain cascading failures and how they spread through synchronous calls
- Combine timeouts, retries with exponential backoff and jitter, and fallbacks
- Implement a circuit breaker and isolate dependencies with bulkheads
Friday night, peak noodle hour. The Loyalty Points service slows to a crawl. Ordering calls it on every order and waits... and waits. Ordering’s threads fill up, so it stops answering the gateway; the gateway’s connections fill up; the whole app is down - because of loyalty points. That’s a cascading failure.
The resilience toolkit (the Distributed Systems track covers retries and idempotency in depth):
| Pattern | Idea |
|---|---|
| Timeout | never wait forever; give up after a budget |
| Retry with backoff and jitter | retry transient failures, waiting longer each time (and randomly), so you don’t pile on a struggling service |
| Circuit breaker | after repeated failures, stop calling for a while and fail fast; then test with a trial call |
| Bulkhead | separate pools (threads, connections) per dependency, so one slow dependency can’t use them all |
| Fallback | a degraded answer instead of an error: cached menu, “points will appear later” |
| Load shedding | under overload, reject some requests quickly to keep serving the rest |
The circuit breaker
A circuit breaker wraps calls to a dependency and has three states:
- Closed (normal): calls go through; failures are counted. After N consecutive failures it opens.
- Open: calls fail immediately without touching the dependency - giving it room to recover and freeing your threads. After a cooldown it goes half-open.
- Half-open: one trial call goes through. Success → closed; failure → open again.
Combined with a fallback, Ordering keeps taking orders while Loyalty recovers, and customers get their points a little later.
1class CircuitBreaker:
2 def __init__(self, threshold, cooldown):
3 self.threshold, self.cooldown = threshold, cooldown
4 self.state, self.failures, self.opened_at = "closed", 0, None
5
6 def call(self, now, succeeds):
7 if self.state == "open":
8 if now - self.opened_at < self.cooldown:
9 return "fast fail (open)"
10 self.state = "half-open"
11 if succeeds:
12 self.state, self.failures = "closed", 0
13 return "ok"
14 self.failures += 1
15 if self.state == "half-open" or self.failures >= self.threshold:
16 self.state, self.opened_at = "open", now
17 return "failed"
18
19breaker = CircuitBreaker(threshold=3, cooldown=10)
20for now, succeeds in [(0, True), (1, False), (2, False), (3, False), (4, True), (9, True), (14, True), (15, True)]:
21 print(f"t={now:>2} {breaker.call(now, succeeds):<16} state={breaker.state}")t= 0 ok state=closed t= 1 failed state=closed t= 2 failed state=closed t= 3 failed state=open t= 4 fast fail (open) state=open t= 9 fast fail (open) state=open t=14 ok state=closed t=15 ok state=closed
Try it
Which pattern helps?
Match each Galactic Noodle Express incident with the pattern that would have helped most.
“A call to Loyalty hung for 10 minutes before failing”
“A brief network blip made 1 in 1,000 payment calls fail; a second try would have worked”
“Loyalty was down for 20 minutes, and every order still waited 2 seconds to time out on it”
“Slow image resizing used up all of the Menu service’s threads, so even cheap menu reads failed”
“The recommendations service is down; the home page showed an error instead of the menu”
Key takeaways
One slow dependency can cascade through synchronous calls and take everything down.
Timeouts on every call; retries only for transient failures, with backoff, jitter and limits.
Circuit breakers fail fast while a dependency recovers; bulkheads keep one dependency from exhausting shared resources.
Fallbacks and load shedding keep the system useful in a degraded state.
Lesson quiz
7 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: simulate microservice patterns in Python
Build small Python simulations of the patterns - routers, sagas, outboxes, circuit breakers, traces - and run them against sample inputs. They run locally in your browser; no servers or containers needed.
Build a circuit breaker
The first input line is threshold cooldown. Each following line is a call time ok|fail (what the dependency would do). Implement the breaker: closed → open after threshold consecutive failures; open calls fail fast until cooldown seconds have passed since it opened; then one half-open trial: success closes it (and resets the count), failure reopens it. Print t=3 failed (closed -> open), t=4 fast fail or t=0 ok - showing a transition in parentheses when the call changes the state (a trial call after the cooldown shows half-open -> closed or half-open -> open). End with calls reaching the dependency: N of M.
- Outage and recovery
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Retries within a deadline
The first input line is base cap max_attempts deadline (seconds). Retry delays grow exponentially: after attempt n (counting from 1) wait min(cap, base * 2**(n-1)) before the next attempt. Each attempt itself takes 1 second. Each following line is a call: the outcomes of its attempts (fail fail ok). For each call, attempt until success, max_attempts is reached, or the next attempt couldn’t finish before the deadline. Print call 1: ok after 3 attempts at 6.0s, call 2: gave up after 4 attempts at 11.0s or ... (deadline).
- Three calls
- Tight deadline
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…