Observability
"How do we know what actually happened?"
Applies To
all components
Why It Happens
A first-principles walkthrough of how you reconstruct what actually happened after a failure. Why a print statement is not observability, how a correlation ID must be carried across every hop, and how structured JSON logs with operation IDs make the timeline queryable in Loki/Datadog after a crash.
How It Works Underneath
Observability is not “more logs.” It is “can you answer what happened to operation X?” A print("transfer failed") cannot be joined to a request, a user, or a provider call. Structured logging carries the correlation: trace_id generated at the edge, propagated in X-Request-ID through every service, emitted as JSON {"trace_id": "req_abc", "payment_id": "pay_123", "duration_ms": 42}. Now a single query — trace_id=req_abc — reconstructs the whole timeline across API, worker, and DB.
Cataloged Failure Modes
untraceable_failuremissing_audit_trail
Code Comparison
# NAIVE: Generic print statements
print(f"transfer {amt}")
# IMPROVED: Structured logging with trace ID
log.info("transfer.completed", trace_id=trace_id, amount=amt, duration_ms=12.4)