Availability
"What happens when a dependency disappears?"
Applies To
databases, payment providers, queues, caches, external services
Why It Happens
A first-principles walkthrough of why a hung dependency can take down your entire fleet. How a thread pool exhausts while waiting on a slow Recs API, why a timeout alone is not enough, and how a circuit breaker (closed→open→half-open) plus bulkheads and fallbacks isolate the failure and fail fast instead of cascading.
How It Works Underneath
Availability is not “the dependency is up.” It is “your system stays up when the dependency is down.” A single await http.get("https://recs.internal") without a timeout holds a worker thread for 30 seconds. Ten concurrent dashboard loads hold ten workers. The pool exhausts, health checks fail, the load balancer marks the node dead, and the outage cascades. The machinery is timeout (bound the wait), circuit breaker (count failures, trip to open, return fallback without calling), and bulkhead (isolate pools so one slow dependency cannot steal all workers).
How to Detect It
Search for any await http without timeout= and without a breaker wrapper. If you cannot find a fallback or get_default_recs() path, the failure will be total, not degraded.
Cataloged Failure Modes
cascadepool_exhaustionhanging_workers
Code Comparison
# NAIVE: Unprotected call blocks thread indefinitely
@app.get("/user/dashboard")
async def get_dashboard(user_id: str):
recs = await http_client.get(f"https://recs.internal/v1/{user_id}")
return {"recs": recs.json()}
# IMPROVED: Circuit breaker + fallback
@app.get("/user/dashboard")
async def get_dashboard(user_id: str):
if breaker.is_open(): return {"recs": await get_default_recs()}
try:
res = await http_client.get(f"https://recs.internal/v1/{user_id}", timeout=1.5)
breaker.record_success()
return {"recs": res.json()}
except Exception:
breaker.record_failure()
return {"recs": await get_default_recs()}