Metadata-Version: 2.4
Name: crowdbench-run
Version: 0.1.4
Summary: Ephemeral, hash-pinned intelligence-benchmark runner for CrowdBench.
Project-URL: Homepage, https://crowdbench.ai
Author-email: Hooch Labs LLC <contact@hoochlabs.com>
Maintainer-email: Hooch Labs LLC <contact@hoochlabs.com>
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Requires-Python: >=3.9
Requires-Dist: httpx
Requires-Dist: platformdirs
Requires-Dist: psutil
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Description-Content-Type: text/markdown

# crowdbench-run

The ephemeral, hash-pinned, open-source intelligence-benchmark runner. Pure Python
(httpx + psutil + platformdirs + stdlib; **never** torch/transformers — HTTP only),
uv-managed, invoked via `uvx` identically on macOS/Linux/Windows.

It probes hardware/engine, normalizes config against the vendored `packages/shared` JSON
artifacts, administers the benchmark item set against a locally-served OpenAI-compatible
endpoint, captures per-item timing forensics, and uploads model **outputs** — scoring is
entirely server-side and answer keys never touch this package. Resumable state lives in
`~/.crowdbench/` (via `platformdirs`).

## Dark-period invocation

No real version is published to PyPI until the repo flip; run from source.

**Bare `crowdbench-run` on a terminal launches the guided wizard** — a zero-file, zero-flag
contribution flow (splash → one-time consent → detect engines → pick a model → pick suites by
category → time budget + a live throughput estimate → reasoning → confirm → run + poll to a final
state). It runs only what this build can administer (core-v1 and code-v1 directly); industry lm-eval suites are
shown with their cost and the exact `crowdbench-run bridge …` command, never run under the wizard's
pretense. On a non-TTY (piped/CI) the bare command prints usage and exits — it never hangs for input.

```
uvx --from ./packages/runner crowdbench-run            # the guided wizard (TTY)
uvx --from ./packages/runner crowdbench-run detect
uvx --from ./packages/runner crowdbench-run inventory
# The flag path (what the wizard drives; agents use it directly). --dataset is now OPTIONAL —
# omit it and the public prompts dataset is auto-fetched from the API into the content-addressed cache.
uvx --from ./packages/runner crowdbench-run core-v1-rc \
    --endpoint http://localhost:8080/v1 --model <model> --agree-tos --yes
```

The RUN flow is **probe → confirm → run → upload**: probe hardware + engine, confirm the plan
(interactive, or `--yes`/`--dry-run` for agents), administer every item capturing output +
per-item timing (capture requirement 12), then POST to `/v1/runs/outputs`, which returns a
pre-issued run URL (status `pending-score` until the server scores). An **interactive** run shows
the same ASCII splash as the wizard (TTY only — never under `--json`, `--yes`, or a pipe), and the
benchmark id is validated **before** the consent screen, so an unsupported id bounces immediately
rather than after you have agreed.

## The default held-out run is a short, server-issued set

A held-out check answers a **server-issued subset** (~20 questions), not the whole pool. The runner
asks the API to deal it (`POST /v1/benchmarks/{id}/challenge`) and administers **exactly** the item
ids it is dealt — it never samples a pool itself, because a client that picks its own "random"
subset picks a favourable one. Scoring accepts only an answered set that matches an issued draw or
the complete pool; anything else parks unscored.

- The challenge is requested **as late as possible** — after consent, after the confirm screen,
  immediately before the first question. A draw is consumed when its answers are scored, so a run
  you abort never burns one.
- `--full` answers the **complete pool** — the opt-in "thorough" run.
- If the API cannot issue one (an older/pre-redeploy deployment), the run **falls back to the
  complete pool and says so**, naming the larger question count. It never falls back to a
  locally-chosen subset.
- A pool no larger than one draw is run whole without asking (the server refuses a draw bigger
  than its pool rather than clamping it).

**These suites are a validity check, not a public score.** They exist because every industry-standard
score is self-reported by the contributor's machine; the held-out questions are the one probe scored
against answers nobody has seen. They place nothing on a leaderboard, and the runner's copy says so
in plain words — user-facing text uses display names ("Reasoning & knowledge check", "Coding check")
and never the internal ids, which stay unchanged in payloads, run URLs, and API paths.

An **interactive** run adds four things (all TTY-only; `--yes`/`--json` surfaces are unchanged):

- **Model picker.** With several loaded models and no `--model`, the wizard's numbered picker
  runs — never a silent first-pick. Non-interactive stays documented behaviour: `--model`, else
  the endpoint's first loaded model.
- **Thinking level, before the confirm screen.** A numbered choice — *model default / off / low /
  medium / high* — that feeds the same `--reasoning-effort` machinery the flag does and is recorded
  identically. `off` sends the upstream-documented `reasoning_effort: none`; support is
  model-dependent, so an engine may ignore or reject it and the run records what the model
  **actually did** (measured from its output), not what was asked. Non-interactive is unchanged:
  the flag, or the model's default.
- **Pre-flight, before the proceed prompt.** What the check is in plain words, how many questions
  against which model, and an **estimated duration range** measured on this machine. The estimate
  is a range, never a single confident number: its low end is the suite's token floor at the
  fastest rate measured, its high end that floor at the slowest rate, multiplied by the suite's
  thinking cost **only when the probe actually observed reasoning** (configuration alone is not
  evidence — an engine may ignore the control). The probe sends the **selected** model's id and
  **discards its first generation**, so a cold weight load never lands in the measured rate. An
  unavailable estimate is reported as unknown, never invented.
- **Honest progress + safe abort.** Progress lines read `question 12/100 · ~1h 40m left` (rolling
  ETA from measured per-item times); raw item ids appear only with `--verbose` (they always ride
  the `--json` payload). `Ctrl-C anytime — partial runs are never uploaded`: SIGINT aborts
  cleanly (exit 130) and uploads **nothing** — a partial suite never scores, so it is never
  submitted, and a started suite is never time-truncated.

`code-v1` runs the identical path: the runner administers the pooled code prompts and uploads
OUTPUTS + timings, and the held-out tests never leave the server (they are the answer key — see
`executor/README.md`). Because code-v1 is scored asynchronously (queue → ephemeral container →
judge) rather than inline, the runner **polls it to a final state** on a longer per-benchmark
window (15 min vs. core-v1's 5); if the window closes first it says so and tells you to report the
run — it never assumes success. code-v1 uses the same sampler defaults as core-v1 (temperature 0,
fixed seed): no in-repo code-v1 spec prescribes a different profile, and the runner does not invent
one. **Consent** is shown once
(state in `~/.crowdbench`); `--agree-tos` attests a human has seen and agreed to the terms (what
uploads; the public CC-BY-NC-SA 4.0 dataset; immutability) and is required to upload non-interactively.

## Commands

| command | what it does |
|---|---|
| *(no arguments, on a TTY)* | the **guided wizard** — zero-file, zero-flag contribution flow |
| `<benchmark-id>` | run one benchmark end-to-end (the flag path); `core-v1` / `core-v1-rc` / `code-v1` / `throughput-std-v1` in this phase |
| `bridge lmeval:<task>:<variant>` | run/parse an allowlisted lm-eval suite → validate → upload |
| `detect` | scan common local serving ports; report apps + loaded models (by route signature) |
| `inventory` | list local models (Ollama / LM Studio / HF cache / user dirs) + benchmark tooling |
| `doctor` | probe + connectivity + state report for support |
| `cache list\|clear` | manage the content-addressed dataset cache |

Key flags (dual-surface parity — every wizard prompt has a flag): `--endpoint`, `--api-base`,
`--dataset` (omit to auto-fetch), `--full` (answer the complete pool instead of the server-issued
subset), `--model`, `--model-path`, `--reps N`, `--seed`,
`--reasoning-effort none|minimal|low|medium|high`, `--reasoning-budget`, `--agree-tos`, `--minutes`,
`--hf-repo`/`--hf-revision`/`--unattributed`, `--endpoint-api-key`
(env `CROWDBENCH_ENDPOINT_API_KEY`; never logged/uploaded/stored), `--dry-run`,
`--no-upload`/`--save`/`--upload-only`, `--json`, `--yes`, `--verbose` (raw item ids on the
progress lines), `--sudo-probes`, `--ports`. Auth uses `CROWDBENCH_API_KEY` (else an anonymous
submitter is minted and cached).

## Trust posture

- **Outputs only.** The runner uploads model outputs + telemetry; it never sees, computes, or
  transmits answer keys. Benchmark item content is inert data — never a tool, never agentic.
- **Dataset integrity.** A dataset's sha256 must match the benchmark's `runner_spec` pin before
  every run (capture requirement 11); a mismatch refuses to run (a tampered set never runs).
- **Vendored artifacts, byte-matched.** `crowdbench_run/_artifacts/*.json` are copied verbatim
  from `packages/shared/artifacts` so both languages read one source; CI fails on drift
  (`scripts/check_artifact_drift.py` — the third leg of the cross-language drift check).
- **Endpoint key privacy.** `--endpoint-api-key` is forwarded as the target endpoint's
  Authorization header only, never logged/uploaded/stored.

Supply-chain audit trail: `packages/runner/BUILD-TIME-CHECKS.md`.

## Publishing (PyPI Trusted Publishing)

Releases go out through `.github/workflows/publish.yml` using **PyPI Trusted Publishing
(OIDC)** — there is **no long-lived PyPI API token** in this repo, ever. GitHub mints a
short-lived OIDC token per run and `pypa/gh-action-pypi-publish` (pinned by full commit SHA)
exchanges it with PyPI. Three human gates stand between a dispatch and a live release:
`workflow_dispatch` only (no push/tag/schedule), a typed `confirm` input that must equal
`publish-to-pypi`, and `environment: pypi` (owner approval before any publish step runs). A
preceding job builds the sdist+wheel, runs `twine check --strict`, install-smokes both
artifacts, and refuses the placeholder `0.0.x` version line.

**Owner one-time setup — verify on pypi.org.** Create the Trusted Publisher for this project
under *Manage project → Publishing* (or *Your projects → Publishing* for a pending publisher).
These **three fields must match** the workflow exactly, or PyPI rejects the OIDC exchange:

| PyPI Trusted-Publisher field | Must be |
|---|---|
| PyPI project (package) name | `crowdbench-run` |
| GitHub repository + workflow filename | `crowdbench-ai/crowdbench-dev`, workflow `publish.yml` |
| Environment name | `pypi` |

(The owner also configures the `pypi` GitHub Environment's protection rule — required
reviewers — so gate #3 actually pauses for approval.) This PR only prepares the workflow and
docs; wiring the Trusted Publisher and the environment protection rule is an owner action on
pypi.org / GitHub, and nothing publishes during the dark period.

## Development

```
cd packages/runner
python -m pip install -e ".[dev]"
python -m pytest -q
```

The test suite runs on ubuntu/macos/windows in CI (a release gate — cross-platform posture).

---

CrowdBench is a service of Hooch Labs LLC. © 2026 Hooch Labs LLC · Apache-2.0.
