Retry Safety
"Is it safe to retry? What retries are allowed?"
Applies To
external calls, webhooks, queue consumers
Why It Happens
A first-principles walkthrough of why retries are inevitable in any distributed system and why an unbounded retry loop is a distributed amplifier. How a tight loop on a 500 turns a single slow dependency into a thundering herd, and how exponential backoff with jitter, a retry budget, and the rule “only retry idempotent operations” make retries safe, bounded, and backpressure-aware.
How It Works Underneath
Retries are inevitable: the network drops, the provider 503s, the user double-clicks. An unbounded retry loop is a distributed amplifier. One slow dependency that returns 500 for 2 seconds, retried immediately by 100 clients, becomes 100× load at the exact moment the dependency is weakest. The fix is three parts: only retry what is idempotent, wait with exponential backoff and jitter so retries spread, and bound the total with a retry budget and a circuit breaker that trips to open after N failures and fails fast without hammering.
Without jitter: 100 clients retry at 1.0s, 2.0s — thundering herd. With jitter: retries scatter.
Common Misconceptions
“More retries = more reliability.” More retries without backoff is a self-inflicted DDoS. Another: “Retry on 400.” 400 is the client’s fault — retrying never fixes a bad request.
Cataloged Failure Modes
retry_stormthundering_herdduplicate_on_retry
Code Comparison
# NAIVE: Tight retry loop
async def send_sms(user_id: str, msg: str):
for _ in range(10):
try: return await sms.send(user_id, msg)
except: time.sleep(0.1)
raise FailedException()
# IMPROVED: Exponential backoff with full jitter
async def send_sms(user_id: str, msg: str, key: str):
for attempt in range(1, 5):
try: return await sms.send(user_id, msg, idempotency_key=f"{key}_{attempt}")
except RetriableException:
delay = random.uniform(0, min(8.0, 0.5 * (2 ** (attempt - 1))))
await asyncio.sleep(delay)
raise FailedException()