Observability Guide

How the platform exposes its internal state.

Metrics

Counters track totals, gauges track current values, and histograms track distributions. All three carry the same labels: service, region, and shard.

Traces

Every request receives a trace id at the edge. The id propagates through every hop and is attached to logs, spans, and any error returned to the caller.

Logs

LevelWhen to useRetention
debuglocal development1 day
inforoutine operation14 days
warnrecoverable anomalies30 days
errorfailed operations90 days

Dashboards

The golden dashboard pairs traffic with latency and errors; every service must appear on it before it takes production traffic.