Metadata-Version: 2.5
Name: ydderd-momentum-cli
Version: 0.8.0
Summary: Momentum customer CLI + experiment-logging SDK — authenticate, upload field data, and report training runs to your workspace.
Project-URL: Homepage, https://momentumbots.io
Project-URL: Repository, https://github.com/momentum-research-labs/momentum
Author: Momentum
License: Apache-2.0
Keywords: cli,ingest,momentum,upload
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Utilities
Requires-Python: >=3.11
Requires-Dist: boto3>=1.43
Requires-Dist: httpx>=0.28
Provides-Extra: dev
Requires-Dist: pytest>=9.1; extra == 'dev'
Requires-Dist: ruff>=0.16; extra == 'dev'
Description-Content-Type: text/markdown

# momentum-cli

The Momentum customer CLI + experiment-logging SDK: authenticate and bulk-upload field data
straight to your workspace's storage bucket, and report training/eval runs from your own compute
into your workspace's experiment tracker.

PyPI distribution: `ydderd-momentum-cli` · Homebrew formula: `momentum-cli` · command: `momentum`.
(The clean `momentum-cli` PyPI name was taken, so the distribution carries the `ydderd-` prefix;
the import package `momentum_cli`, the `momentum` command, and the brew name are unaffected.)

## Why this is a separate package

The CLI talks to the Momentum API purely over HTTP, and to Momentum's S3 storage with
short-lived credentials the API mints per upload. It shares **no Python code** with the backend,
so it ships with a tiny dependency set — `httpx` + `boto3` — instead of the full server stack
(torch, opencv, fastapi, …). That keeps the install small and avoids shipping the backend's AGPL
detector to customers.

## Install

```bash
brew install momentum-research-labs/momentum/momentum-cli
# or:
pipx install ydderd-momentum-cli
momentum --help
```

## Use it with a coding agent

The CLI ships an agent skill — a description of these commands written for Claude Code, Cursor or
Codex. It lives inside the package, so it always matches the version you have installed:

```bash
momentum skill install     # writes .claude/skills/momentum/SKILL.md
momentum skill print       # or just read it
```

Re-run with `--force` after upgrading the CLI. `--dir` writes somewhere other than
`.claude/skills/momentum/`.

## Usage

```bash
momentum auth login                       # opens a browser for workspace approval
momentum auth whoami                      # confirm tenant
momentum datasets list                    # list datasets (id, name, kind) to pick a target
momentum upload ./your-data --dataset "Evaluation rollouts"  # bulk upload, lands in that Dataset
momentum datasets assign <session_id> --dataset "Evaluation rollouts"  # assign an upload later
momentum ingest status                    # ingest ledger stats
momentum ingest scan                      # submit a scan and print its job_id
momentum metrics dataset <dataset_id>     # each measure's spread, and the Dataset's thresholds
momentum metrics operators <dataset_id>   # per-operator error rates and driving-style deviations
momentum analysis status <dataset_id>     # whether analysis is current: ready, stale, queued, running, failed
momentum analysis run <dataset_id>        # start the ordinary analysis run; a no-op while one is queued or running
momentum findings list <dataset_id>       # the Dataset's findings and whether each is resolved
momentum findings resolve <dataset_id> --key <key> --note "…"  # resolve one finding; reopen with `findings reopen`
momentum metrics episodes <dataset_id> --limit 200  # a page of Episodes with metric values; --cursor for the next
momentum metrics episodes <dataset_id> --operator <op_id> --tag <tag> --title <text> --has-depth  # narrowed
momentum episodes get <episode_id> [<episode_id> ...] --media  # named Episodes: measures, stall, app link; --media adds signed video URLs
momentum benchmarks create --name "RoboArena-DROID" --spec pins.json  # mint a benchmark, print its id
momentum benchmarks list --latest         # benchmark versions (one per family) to pick a pin
momentum registry list                    # this workspace's registry instances
momentum registry instantiate --class-ref ood-detector@2 --bindings bindings.json
momentum eval run --benchmark pebk-roboarena@3 --mode smoke  # queue the platform CPU smoke path
momentum eval submit --model hf://lab/pi05-fast --benchmark pebk-roboarena@3  # open an eval run, print URL
momentum trials log --policy hf://lab/pi05-fast --n 20 --successes 14 --calibration-set pcsk-…
momentum trials log --csv trials.csv      # bulk floor tallies (per-row partial success)
momentum secrets set lab-bucket           # store a secret (value from stdin); prints its creds_ref
momentum secrets list                     # secret names + configured (never values)
momentum gpu credentials                  # Momentum team: an hour of AWS access for GPU queue work
```

Process an upload once it is assigned. A large upload is processed in parts on several workers;
`--wait` prints how many parts have finished and exits non-zero if the run failed.

```bash
momentum uploads process 616800d20685 --embodiment arm-v1 --revision 1 --dataset-type cloudchef_trajectory_v1 --wait
momentum uploads status 616800d20685      # latest run, and each failed part with where it stopped
```

Running `uploads process` again after a failure retries only the failed parts.

`momentum eval run --mode smoke` is the platform's deployed AWS CPU smoke path. It requires either an exact
`peb-…` benchmark id or a `pebk-…@positive-version` pin, plus the workspace `evaluation`
entitlement, then prints the execution receipt and URL. It does not load a policy or simulator,
produce scores/OOD, rank policies, or change inference deployment. Use `momentum eval submit` for
an external process that reports a real policy evaluation through the SDK.

`--mode world-model` uses the same admission as the Eval Runs launch form. The exact benchmark
pins the world model; `--policy` must identify a supported registered runtime, not just a model
identity. Verify first without allocating compute, then submit with an explicit retry identity:

```bash
momentum eval run --mode world-model --benchmark pebk-demo@1 --scenario scn-example \
  --policy gs://openpi-assets/checkpoints/pi05_base --verify-only
momentum eval run --mode world-model --benchmark pebk-demo@1 --scenario scn-example \
  --policy gs://openpi-assets/checkpoints/pi05_base --idempotency-key my-diagnostic-001
```

Reuse the same key after a lost response; a new key is a new launch intent. Submission repeats
input verification and prints the execution record. This diagnostic runs two feedback windows;
it does not train weights, score task success, or establish physical fidelity. An operator must
configure the frozen runtime image, private seed delivery and shared compute pool.

`momentum benchmarks create` mints version 1 and prints its id — which is what `eval submit
--benchmark` takes. Until it existed the only way to get that id was a browser form, so a headless
run could not start from nothing. `--spec` carries the nested pins (`target`, `metrics`,
`evaluator_stack`, `run_config`) as a JSON object: inline, a path, or `-` to read stdin. Pass
`--not-primary` to mint a version without making it the family's primary.

`momentum registry instantiate` pins a global class into your workspace. `--class-ref` takes an
`rgc-…` id, a `<class_key>@<version>` pin, or a bare family key for the latest; `--bindings` takes
the tenant particulars as JSON, by the same three routes as `--spec`.

`upload` always writes to your workspace's one raw prefix — there's no target to choose. Whether
what you uploaded is raw drone video (needs extraction) or already-extracted frames is classified
server-side once it lands, not by the client beforehand.

Bytes never pass through the API. Each upload session gets S3 credentials that are valid only
for that session's prefix; the CLI PUTs files straight to the bucket with them, then confirms
what landed. Credentials are issued a few hours at a time and the CLI renews them on its own, so
one `momentum upload` can run for days — a multi-terabyte dataset needs one command, not a
babysitter. Progress prints every few seconds (`38.2/512.0 GB · 185 MB/s · 1,204/9,412 files`).

A file that fails is retried (five attempts, backing off) and does not stop the run; the ones
that still fail are listed at the end together with the command that finishes the job:
`momentum upload <dir> --resume <session_id>` sends only what is missing — a file already in
place is checked and skipped, never sent again. Ctrl-C lets the files in flight finish, then
prints the same resume command. A session stays open for 7 days after its last activity (a
confirmed file or a credential renewal); after that, start a new upload.

`--dataset NAME` (or `--dataset-id ID`) checks the Dataset exists before any bytes move and
assigns the upload to it the moment it completes — the same assignment `momentum datasets
assign` makes, for uploads that were left to be triaged in the app.

`momentum ingest scan` and `momentum upload --scan` submit the scan as a background job and print
its `job_id`; neither command waits for the scan itself to finish.

For headless/CI use, skip the browser with a token minted by any provisioned workspace user:
`momentum auth login --token <fw_cli_…>`.

If browser approval fails or times out, the CLI exits with a retry message instead of a traceback.
Run `momentum auth login` again; if you switched workspaces in the browser, refresh the app and
approve from the target workspace.

Config is stored at `~/.momentum/config.json` (an existing `~/.flywheel/config.json` is copied
over once on first use). Auth precedence: `MOMENTUM_CLI_TOKEN` env > config file.

## Experiment-logging SDK

Training and eval runs executed on your own compute (Modal, Brev, a lab box) report themselves
into your workspace's experiment tracker — W&B-style, and safe to leave in production training
code (a logging failure never raises into the train):

```python
import momentum_cli as momentum

run = momentum.init(name="my_sft_run", tags=["sft"], config={"iters": 800, "lr": 2e-4},
                    provider="modal")
run.log({"train/loss": 0.42}, step=100)
run.finish(status="succeeded", checkpoint_ref="s3://lab-bucket/ckpt")

# later — scoring results and billed cost arrive after the train, so annotation
# works on finished runs:
momentum.annotate(run.id, results={"auroc": {"value": 0.61, "ci": [0.55, 0.67]}})
```

### GPU jobs from a Metaflow flow (Momentum team)

A Metaflow step can run its GPU work on Momentum's GPU queue. `run_gpu_job` submits the job, records
it as an experiment, waits for it (including for an admin to approve a paid plane), logs the loss
curve the job wrote, and returns where the outputs landed:

```python
from metaflow import FlowSpec, current, step
from momentum_cli.gpu import run_gpu_job

class TrainFlow(FlowSpec):
    @step
    def train(self):
        self.gpu = run_gpu_job(job_spec, current=current, params={"lr": self.lr})
        self.next(self.end)
```

`eval "$(momentum gpu credentials)"` loads an hour of AWS access for pushing the job's image and
reading its outputs. Only a key for the Momentum workspace can do either. A worked flow is in
`examples/gpu_training_flow.py`.

### Eval runs (policy context — the CI-integration path)

An eval process (a lab rig, Modal, the robot) creates one Episode for each execution, then attaches
the evaluation result to that Episode. Results share the training SDK's never-raise, heartbeat, and
reattach behavior; they buffer and flush in idempotent batches:

```python
import momentum_cli as momentum

ev = momentum.eval_run(benchmark="pebk-roboarena@3", model="hf://lab/pi05-fast", seeds=3)
episode = momentum.episodes.create(video={"camera": "rollout-0.mp4"})
ev.log_result(episode_id=episode.id, scenario="scn_pick", seed=0, status="success",
              scorer={"success": True, "task_progress": 1.0}, latency_p50=61.0)
episode = momentum.episodes.create(video={"camera": "rollout-1.mp4"})
ev.log_result(episode_id=episode.id, scenario="scn_pick", seed=1, status="fail",
              scorer={"success": False})
ev.finish()                                   # flushes any buffered results first
ev.annotate(results={"headline": {"value": 0.5, "ci": [0.31, 0.69]}})   # post-hoc scoring
```

`eval_run()` prints the run URL on create; `eval_run(run_id=…)` (or `MOMENTUM_EVAL_RUN_ID`) reattaches
after a preemption. Runs land in the UI under **Eval runs**.

### Real trials (floor tallies → calibration audit)

Report real-robot trials of a policy; landing trials that ground a calibration set recomputes that world
model's τ/ρ trust:

```python
momentum.real_trials.log(policy="hf://lab/pi05-fast", scenario="scn_pick",
                         n=20, successes=14, operator="alice", calibration_set="pcsk-…")

report = momentum.real_trials.log_csv("trials.csv")   # a path or raw CSV text; per-row partial success
print(report["accepted"], report["rejected"])
```

### Episode ingestion and secrets

Episode creation uses the same upload-session workflow as every other ingest path. Upload the source
dataset unchanged, then assign the returned session to its Dataset. Normalization creates database-owned
Episode identities and the application-readable artifacts:

```bash
momentum upload ./session_042 --dataset "Training demonstrations"
# or, to triage the upload in the app first:
momentum upload ./session_042
momentum datasets assign <session_id> --dataset "Training demonstrations"
```

Read one cursor page of the Dataset's normalized Episodes to verify an ingest or discover IDs for
curation and snapshots:

```python
page = momentum.episodes.list(dataset_id="<dataset_id>", limit=100)
while page:
    for episode in page["items"]:
        print(episode["id"], episode["task_label"], episode["outcome"])
    cursor = page["page_info"]["end_cursor"]
    if cursor is None:
        break
    page = momentum.episodes.list(
        dataset_id="<dataset_id>", limit=100, cursor=cursor
    )
```

`limit` defaults to 50 and is capped at 200. The call returns one page with `items`, `page_info`, and
`total_count`; it never loads the tenant's complete Episode corpus implicitly.
Named secrets remain available for supported external-service configuration; values are encrypted at
rest and never returned by a read.

Auth: `MOMENTUM_API_KEY` env (a `fw_cli_…` token — inject it as a secret in your training
environment), falling back to the token saved by `momentum auth login`. `MOMENTUM_API_URL`
overrides the API endpoint. `with momentum.init(...) as run:` (and `momentum.eval_run(...)`) marks the
run failed (with the exception) if the block raises. Runs land in the workspace UI under
**Experiments** / **Eval runs**.

Release/consumption mechanics (PyPI, git-ref installs, versioning): see `PUBLISHING.md`.

## Developer notes

These knobs exist for Momentum developers and are intentionally hidden from customer-facing
help and docs:

- **`--api-url <url>` on `momentum auth login`** — persist a non-production API base URL to the
  config (e.g. a local API). Hidden via `argparse.SUPPRESS`.
- **`MOMENTUM_API_URL` env** — override the API base per-invocation. Takes precedence over the
  config file.

Precedence for the API base URL: `MOMENTUM_API_URL` env > `api_url` in config > default
(`https://api-aws.momentumbots.io/api`, the hosted production API on AWS).

Point the CLI at a local backend during development:

```bash
MOMENTUM_API_URL=http://localhost:8000 momentum auth whoami
# or persist it:
momentum auth login --token <fw_cli_…> --api-url http://localhost:8000
```

### Local development

```bash
cd cli
uv sync
uv run momentum --help
uv run pytest
```

## Releasing (PyPI + Homebrew)

PyPI is the source of truth; the Homebrew formula wraps the published PyPI sdist.

### 1. Publish to PyPI — via GitHub Actions (Trusted Publishing, no token)

The `.github/workflows/publish-cli.yml` workflow builds and publishes over OIDC. Cut a release
by pushing a namespaced tag from the monorepo default branch:

```bash
git tag cli-v0.1.0 && git push origin cli-v0.1.0
```

The PyPI project is `ydderd-momentum-cli`, published from `momentum-research-labs/momentum` via the `pypi`
environment. (First publish activates the "pending" Trusted Publisher and creates the project.)

### 2. Update the Homebrew tap formula

After the PyPI release exists, point `release.sh` at your tap checkout — with `SKIP_PUBLISH=1`
it skips the upload, fetches the published sdist's `url`/`sha256`, writes an explicit formula
`version`, refuses placeholder formula values, and regenerates Python `resource` blocks:

```bash
SKIP_PUBLISH=1 \
FORMULA_PATH=/path/to/homebrew-momentum/Formula/momentum-cli.rb \
  cli/scripts/release.sh
```

Then commit + push the tap. Customers install with:

```bash
brew install momentum-research-labs/momentum/momentum-cli
```

> `release.sh` can also publish to PyPI itself (`UV_PUBLISH_TOKEN=pypi-… cli/scripts/release.sh`)
> if you prefer a token-based local release over the GitHub Action.

Bumping a release: change `version` in `pyproject.toml` and `src/momentum_cli/__init__.py`,
push a new `cli-v*` tag, then re-run step 2. `release.sh` refuses to continue if those versions
drift.
