Failures
Failures
GitHub
Home Docs Recovery
MEDIUM

Recovery

"How does system return to valid state?"

Invariant: Every failure has a recovery path to a valid state.

Applies To

queues, payments, enrollments, file uploads

Why It Happens

A first-principles walkthrough of how a system returns to a valid state after a poison message or a crash before ack. Why an at-least-once queue without an idempotent consumer stays broken, and how a dead-letter queue plus a reconciliation sweeper are the two recovery paths that prevent pending-forever.

How It Works Underneath

Recovery is the second half of every failure. An at-least-once queue that crashes before ack will redeliver. Without an idempotent consumer, the certificate is sent twice. Without a dead-letter queue, a poison message that always throws JSON parse error will be retried forever, stalling the partition. The machinery is two paths: the happy path acks after durable processing and deduplicates via msg_id UNIQUE, the sad path moves the poison to a DLQ after N attempts and a reconciler sweeps pending rows by asking the provider.

Cataloged Failure Modes

Code Comparison

Language:
Fragile (AI Happy Path) — Python recovery_fragile.py
# NAIVE: Poison message crashes consumer
async def consume():
    msg = await q.pop()
    data = json.loads(msg.body)
    await q.ack(msg)
Resilient (Failures Verified) — Python
recovery_safe.py
# IMPROVED: Dead Letter Queue routing on failure
async def consume():
    msg = await q.pop()
    try:
        data = json.loads(msg.body)
        await q.ack(msg)
    except Exception as e:
        await dlq.push(msg, error=str(e))
        await q.ack(msg)

Mitigation Patterns