Failures
Failures
GitHub
Home Docs Observability
MEDIUM

Observability

"How do we know what actually happened?"

Invariant: Every operation must be traceable to its outcome.

Applies To

all components

Why It Happens

A first-principles walkthrough of how you reconstruct what actually happened after a failure. Why a print statement is not observability, how a correlation ID must be carried across every hop, and how structured JSON logs with operation IDs make the timeline queryable in Loki/Datadog after a crash.

How It Works Underneath

Observability is not “more logs.” It is “can you answer what happened to operation X?” A print("transfer failed") cannot be joined to a request, a user, or a provider call. Structured logging carries the correlation: trace_id generated at the edge, propagated in X-Request-ID through every service, emitted as JSON {"trace_id": "req_abc", "payment_id": "pay_123", "duration_ms": 42}. Now a single query — trace_id=req_abc — reconstructs the whole timeline across API, worker, and DB.

Cataloged Failure Modes

Code Comparison

Language:
Fragile (AI Happy Path) — Python observability_fragile.py
# NAIVE: Generic print statements
print(f"transfer {amt}")
Resilient (Failures Verified) — Python
observability_safe.py
# IMPROVED: Structured logging with trace ID
log.info("transfer.completed", trace_id=trace_id, amount=amt, duration_ms=12.4)

Mitigation Patterns