Metadata-Version: 2.5
Name: ydderd-momentum-sdk
Version: 0.1.1
Summary: Momentum edge SDK — wrap a black-box model, score real-time uncertainty, flag anomalous events, and report telemetry.
Project-URL: Homepage, https://withflywheel.com
Project-URL: Repository, https://github.com/ydderd/momentum
Author: Momentum
License: Apache-2.0
Keywords: momentum,ood,out-of-distribution,robotics,telemetry,uncertainty
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: torch
Provides-Extra: capture
Requires-Dist: mcap>=1.2; extra == 'capture'
Provides-Extra: dev
Requires-Dist: mcap>=1.2; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Description-Content-Type: text/markdown

# momentum-sdk

The Momentum edge SDK: wrap a black-box model, score real-time uncertainty in your inference
loop, flag anomalous events for review, and (later) report telemetry — one package, one install.
Named for the whole edge-runtime surface rather than "momentum-ood", since OOD detection is the
first capability, not the only one. Free wedge into the Momentum platform — flagged events route
into curated train/eval sets.

PyPI distribution: `ydderd-momentum-sdk` · import package: `momentum_sdk`.

## Why this is a separate package from the CLI

`ydderd-momentum-cli` talks to the Momentum API over HTTP/S3 and ships with a tiny dependency
set. This SDK wraps your own `torch.nn.Module` in your real-time (10–30Hz) loop, so it has a
genuine `torch` dependency — kept unpinned since torch is your install, not ours to pin. See
[PUBLISHING.md](PUBLISHING.md) for the release path.

## What's here

- **Model wrapper** — wrap a black-box `nn.Module`: fingerprinting, activation capture, and a
  measured compute-overhead budget that keeps scoring off the critical path.
- **State capture** — register frame / proprioception / action sources into per-source ring
  buffers behind a single `capture_fidelity` lever; `snapshot()` materializes the rolling
  `[t−Δ, t]` window.
- **Scorer interface** — a pluggable `Scorer` with conformal thresholds, `MaxScoreCombiner`,
  and per-method `ScorerConfig`; see the [trade-off table](SCORERS.md).
- **Calibrated scorers** — relative Mahalanobis (`rmds`) for classifier/detector heads and
  logpZO (`logpzo`) for diffusion/flow policies. A model matching neither is reported, not
  scored with an uncalibrated fallback.
- **`init` walkthrough** — one call from a live model to a calibrated, already-scoring detector.
- **Control-loop telemetry** — inference latency, teleop latency, handoff rates and autonomy
  fractions, measured by the same wrapper (below). No scorer or calibration required.
- **Metadata telemetry** — opt-out, install-keyed; raw data never rides telemetry.
- **Automated setup** — a coding-agent skill + `momentum-analyze` static scan (below).

Snapshot upload to the platform is not part of this package yet.

## Automated setup

The SDK ships a coding-agent skill (`momentum_sdk/skills/momentum-setup/SKILL.md`) and the
static scan it drives:

```bash
momentum-analyze <repo-or-package>       # locate the model, its loop, state sources, and
                                         # existing failure sinks (add --json for structure)
```

Point your coding agent at the skill and it proposes a reviewable diff — wrap the model with
`init`, register the state sources, and add flags beside your existing logging/Sentry/ROS
sinks. It proposes; it never edits production code without showing the diff first.

## Usage

Three primitives — **init, calibrate, run**:

```python
import momentum_sdk as momentum

detector = momentum.init(model, on_flag=lambda score: alert(score))  # 1. init
detector.calibrate(nominal_pairs)                                    # 2. calibrate
output = detector.run(x)                                             # 3. run (also detector(x))
```

`init` inspects the model and picks its calibrated scorer (`rmds` for classifier/detector
heads, `logpzo` for diffusion/flow policies). `calibrate` fits the flagging threshold on
`nominal_pairs` — an iterable of `(activations, output)` pairs from the model behaving
normally (for `logpzo`, successful rollouts only). `run` scores every call and flags once
calibrated (running before `calibrate` is inert — it captures state but doesn't score or
flag until calibrated). A model
matching neither scorer raises `UnsupportedLeafError` at `init` after reporting its
fingerprint — there is no uncalibrated fallback. `detector.summary()` recaps what was
detected, picked, and calibrated.

Pass `nominal_data=` to `init` to fold calibrate into the init call as a one-shot. Or wire
the wrapper directly for the advanced path (`OODDetector(model, scorer=…)` or `score_fn=…`):

```python
detector = momentum.OODDetector(model, target_hz=20.0)

output = detector(inputs)     # same call signature as model(inputs)

detector.fingerprint          # ModelFingerprint: framework, head type, backbone, deployment target
detector.overhead             # {"mean_ms", "p95_ms", "budget_ms", "dropped_scores"}
```

### Robot identity

Set `robot_id` when a detector observes one known physical fleet unit:

```python
detector = momentum.init(model, robot_id="arm-07")

# The direct constructor exposes the same option.
detector = momentum.OODDetector(model, robot_id="arm-07")
```

`robot_id` is a customer-defined, tenant-scoped reference. The SDK does not infer it from the
hostname, hardware, installation, deployment, or process run. One detector observes one robot
stream; create separate detector instances when one process observes multiple robots. Every event
inherits the detector's value before spooling. Omitting it leaves events valid but unattributed.

Robot identity is sent only with captured events. Anonymous SDK telemetry and object-storage paths
remain unchanged. The platform exposes it on event records and supports fleet-history filtering with
`GET /events?robot_id=arm-07`.

### Compute-overhead budget

The wrapper measures its own added latency (wrap + hook capture + scoring) against a budget
relative to the loop period — `budget_ms = (1000 / target_hz) * overhead_budget_fraction`
(10% by default). Scoring runs off the critical path on a single background worker; if it's
still busy when the next call arrives, that call's scoring is dropped rather than queued, so a
slow scorer never stalls inference.

```python
detector = momentum.OODDetector(
    model,
    target_hz=20.0,
    on_overhead_exceeded=lambda stats: logger.warning("OOD overhead over budget: %s", stats),
)
```

### State capture

Two tiers feed a rolling `[t−Δ, t]` window of per-source ring buffers (timestamp +
step-index aligned, bounded, non-blocking):

- **Passive** — the wrapper steps the buffer once per call and records the cheap
  model-boundary channel (`model.output_norm`) automatically.
- **Registered** — you declare the sources the SDK can't see from inside a black-box module:

```python
# pull-style: polled once per wrapped call; a failing getter is counted, never raised
detector.capture.register("frame.wrist", getter=lambda: camera.latest, kind="frame", rate_hz=30)

# push-style: from wherever the data already flows (vector values expand to name.0…name.n-1)
detector.record("proprio.joints", robot.joint_positions)
detector.record("action", action_vector)

snap = detector.snapshot()   # channels / frames / sources, in the platform event-bundle shape
```

Everything sits behind one `capture_fidelity` lever — `off` / `light` (scalars only, 5s) /
`standard` (+frames at 1-in-4 stride, 10s) / `full` (every frame, 30s) — with `capture_window_s`
as the single per-deployment override. Uploading a snapshot on a flag, and edge durability via
a local on-device log, are not part of this package yet.

### Control-loop telemetry — latency, handoffs, and autonomy

Wrapping the model puts the SDK on the one code path that runs every control step, so the same
wrapper measures what that loop costs and who is driving it. All of it works on any detector —
scored or bare — and needs no calibration.

**Inference latency** is measured automatically on every wrapped call. It's kept apart from the
wrapper's own `overhead`, so a slow model never reads as SDK cost, and vice versa:

```python
detector.inference_latency   # {"mean_ms", "p50_ms", "p95_ms", "p99_ms", "max_ms", "samples", "count"}
```

**Teleop** is the half the SDK can't see on its own — it lives in your stack, so you mark it.
A takeover is one `teleop()` block; one command from the teleop stack is one `teleop_command()`:

```python
with detector.teleop("operator took over — bin was empty"):   # the takeover: counted + timed
    while operator.engaged:
        with detector.teleop_command():                       # one command: its round-trip latency
            robot.apply(operator.command())
```

`begin_teleop()` / `end_teleop()` are the same thing without a block, and
`record_control_source("teleop")` drives the ledger from a stack that already tracks who's
driving. If your teleop stack already times its own round trip, pass it straight in with
`record_teleop_latency_ms(ms)`.

Read it all back at any time:

```python
detector.metrics()      # {"inference": …, "teleop": …, "control": …, "overhead": …}
```

The `control` half answers the operational questions:

| metric | what it answers |
|---|---|
| `teleop_time_fraction` / `autonomy_time_fraction` | how much of the driven time a person did |
| `teleop_step_fraction` | the same split in control steps, the live analog of the `human_fraction` characterization computes per frame |
| `handoffs_to_teleop` / `handoffs_per_hour` | how often someone has to step in |
| `mean_teleop_s` / `longest_teleop_s` | how long they stay in once they do |
| `mean_autonomous_s_between_handoffs` | how long it runs before needing help |
| `shadow_inferences` | inference spent while a human was driving anyway |
| `handoffs_after_flag` / `mean_flag_lead_s` | how many takeovers the SDK flagged first, and how much warning it gave |

A control source is an open string, and the words that mean *a person drove* — `"teleop"`,
`"human"` — are the same ones dataset characterization matches on, so the live autonomy number
and the one computed later over the uploaded episodes are the same measure. `"none"` / `"idle"`
name a gap rather than a driver, so `policy → none → teleop` is one handoff, not two. A stack
with its own vocabulary passes it in: `ControlLedger(teleop_sources={"operator"})`.

A takeover also captures an event carrying the `[t−Δ, t]` window that led up to it — a person
grabbing the controls is a human labelling the moment autonomy stopped being good enough, and
those seconds are the clip worth training on. Pass `capture_teleop_events=False` to count
takeovers without materializing a window for each.

One takeover captures exactly one window, whichever call declared it: re-declaring teleop while
a person is already driving does nothing, and a nested `teleop()` block (a helper decorated with
`@detector.teleop()`, called from inside one) leaves control with the operator on the way out
rather than handing it back early. So `record_control_source(src)` is safe to call every control
step — only a real change of driver does anything.

Two metadata-only telemetry events carry this: `control-handoff` on each driver change, and a
`runtime-metrics` roll-up at `close()` (or whenever you call `report_metrics()`). Latencies,
counts and fractions ride; the *classified* driver rides, never the raw control-source word or
your `reason` text — those stay local, on the captured event.

The ledger works standalone, for a stack whose inference the wrapper can't reach:

```python
from momentum_sdk import ControlLedger

control = ControlLedger()
control.record_step("teleop")          # or set_source / begin_teleop / end_teleop
control.stats()
```

### Error-reporting mode (when the model has no calibrated leaf)

The SDK never emits an uncalibrated OOD score — so if a model isn't a classifier/detector head
(`rmds`) or a diffusion/flow policy (`logpzo`), `init(model)` raises `UnsupportedLeafError`
rather than guess. That isn't a dead end: `init(model, mode="error_reporting")` returns a **bare**
detector that still adds value, working like a Sentry SDK for the robot loop — you capture events
and exceptions manually, each carrying the same `[t−Δ, t]` capture window a scored flag would.
It just can't catch the failures that *don't* throw; OOD scoring turns on later, once the model's
leaf ships.

```python
detector = momentum.init(model, mode="error_reporting")   # no scorer, no calibration step

detector.flag("gripper slipped", level="warning")         # manual event
detector.capture_message("unexpected object in bin")       # Sentry-parity alias
try:
    plan = policy(obs)
except Exception as err:
    detector.capture_exception(err)                        # captures the event + window, re-raise is yours

with detector.capture_errors():                            # captures anything raised here, then re-raises
    step(env)
```

Events go to the `on_event` callback you pass; `before_send` is an opt-in redaction hook, and
nothing leaves the machine unless you opt in (network upload isn't in this package yet). The same
capture surface (`flag` / `capture_message` / `capture_exception` / `capture_errors`) is available
on a scored detector too — error reporting is additive, not a separate mode you're locked into.
