Recovery
"How does system return to valid state?"
Applies To
queues, payments, enrollments, file uploads
Why It Happens
A first-principles walkthrough of how a system returns to a valid state after a poison message or a crash before ack. Why an at-least-once queue without an idempotent consumer stays broken, and how a dead-letter queue plus a reconciliation sweeper are the two recovery paths that prevent pending-forever.
How It Works Underneath
Recovery is the second half of every failure. An at-least-once queue that crashes before ack will redeliver. Without an idempotent consumer, the certificate is sent twice. Without a dead-letter queue, a poison message that always throws JSON parse error will be retried forever, stalling the partition. The machinery is two paths: the happy path acks after durable processing and deduplicates via msg_id UNIQUE, the sad path moves the poison to a DLQ after N attempts and a reconciler sweeps pending rows by asking the provider.
Cataloged Failure Modes
poison_messageredelivery_without_idempotencypending_forever
Code Comparison
# NAIVE: Poison message crashes consumer
async def consume():
msg = await q.pop()
data = json.loads(msg.body)
await q.ack(msg)
# IMPROVED: Dead Letter Queue routing on failure
async def consume():
msg = await q.pop()
try:
data = json.loads(msg.body)
await q.ack(msg)
except Exception as e:
await dlq.push(msg, error=str(e))
await q.ack(msg)