Metadata-Version: 2.4
Name: crowdbench-run
Version: 0.1.6
Summary: Ephemeral, hash-pinned intelligence-benchmark runner for CrowdBench.
Project-URL: Homepage, https://crowdbench.ai
Author-email: Hooch Labs LLC <contact@hoochlabs.com>
Maintainer-email: Hooch Labs LLC <contact@hoochlabs.com>
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Requires-Python: >=3.9
Requires-Dist: httpx
Requires-Dist: platformdirs
Requires-Dist: psutil
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Description-Content-Type: text/markdown

# crowdbench-run

The ephemeral, hash-pinned, open-source intelligence-benchmark runner. Pure Python
(httpx + psutil + platformdirs + stdlib; **never** torch/transformers — HTTP only),
uv-managed, invoked via `uvx` identically on macOS/Linux/Windows.

It probes hardware/engine, normalizes config against the vendored `packages/shared` JSON
artifacts, administers the benchmark item set against a locally-served OpenAI-compatible
endpoint, captures per-item timing forensics, and uploads model **outputs** — scoring is
entirely server-side and answer keys never touch this package. Resumable state lives in
`~/.crowdbench/` (via `platformdirs`).

## Dark-period invocation

No real version is published to PyPI until the repo flip; run from source.

**Bare `crowdbench-run` on a terminal launches the guided wizard** — a zero-file, zero-flag
contribution flow (splash → one-time consent → detect engines → pick a model → measure speed →
suites grouped by category → time budget + a live throughput estimate → reasoning → confirm → run
+ poll to a final state → wrap-up). It runs only what this build can administer (core-v1 and
code-v1 directly); industry lm-eval suites appear **under their own category** with their cost and
the exact `crowdbench-run bridge …` command, never run under the wizard's pretense. The coding
category says out loud that it has no industry suite by design: every standard one scores by
executing model-generated code on *your* machine, which CrowdBench never does. On a non-TTY
(piped/CI) the bare command prints usage and exits — it never hangs for input.

Two things the wizard settles before it asks you anything else:

- **Are the questions actually published?** Each directly-runnable suite's item set is probed while
  the catalog is built (the same cached fetch the run itself uses, so the run pays nothing extra).
  A suite the server has not published is greyed with the reason and excluded — rather than failing
  after you have answered every prompt. A probe that merely *fails* (timeout, 5xx, offline) never
  hides a suite; only a definitive "not published" does.
- **Does the plan still fit the budget you gave?** If the estimate on the confirm screen exceeds the
  time budget from earlier in the session, it says so in one loud line before the proceed prompt.

The times the wizard quotes follow the thinking level you pick. A suite's estimate is a
non-thinking **floor**; choosing *low/medium/high* multiplies it by that suite's thinking cost,
choosing **off** leaves it at the floor, and leaving it at *model default* shows a **range** —
`~88s–24m, depends how much the model chooses to think` — because a model that thinks by default
lands near the top and one that doesn't lands at the floor, and nothing has decided which yet.
Budget filtering always uses the floor: a check is excluded for time it will certainly take, never
for time the model may or may not spend.

```
uvx --from ./packages/runner crowdbench-run            # the guided wizard (TTY)
uvx --from ./packages/runner crowdbench-run detect
uvx --from ./packages/runner crowdbench-run inventory
# The flag path (what the wizard drives; agents use it directly). --dataset is now OPTIONAL —
# omit it and the public prompts dataset is auto-fetched from the API into the content-addressed cache.
uvx --from ./packages/runner crowdbench-run core-v1-rc \
    --endpoint http://localhost:8080/v1 --model <model> --agree-tos --yes
```

The RUN flow is **probe → confirm → run → upload**: probe hardware + engine, confirm the plan
(interactive, or `--yes`/`--dry-run` for agents), administer every item capturing output +
per-item timing (capture requirement 12), then POST to `/v1/runs/outputs`, which returns a
pre-issued run URL (status `pending-score` until the server scores). An **interactive** run shows
the same ASCII splash as the wizard (TTY only — never under `--json`, `--yes`, or a pipe), and the
benchmark id is validated **before** the consent screen, so an unsupported id bounces immediately
rather than after you have agreed.

## The default held-out run is a short, server-issued set

A held-out check answers a **server-issued subset** (~20 questions), not the whole pool. The runner
asks the API to deal it (`POST /v1/benchmarks/{id}/challenge`) and administers **exactly** the item
ids it is dealt — it never samples a pool itself, because a client that picks its own "random"
subset picks a favourable one. Scoring accepts only an answered set that matches an issued draw or
the complete pool; anything else parks unscored.

- The challenge is requested **as late as possible** — after consent, after the confirm screen,
  immediately before the first question. A draw is consumed when its answers are scored, so a run
  you abort never burns one.
- `--full` answers the **complete pool** — the opt-in "thorough" run.
- If the API cannot issue one (an older/pre-redeploy deployment), the run **falls back to the
  complete pool and says so**, naming the larger question count. It never falls back to a
  locally-chosen subset.
- A pool no larger than one draw is run whole without asking (the server refuses a draw bigger
  than its pool rather than clamping it).

**These suites are a validity check, not a public score.** They exist because every industry-standard
score is self-reported by the contributor's machine; the held-out questions are the one probe scored
against answers nobody has seen. They place nothing on a leaderboard, and the runner's copy says so
in plain words — user-facing text uses display names ("Reasoning & knowledge check", "Coding check")
and never the internal ids, which stay unchanged in payloads, run URLs, and API paths.

An **interactive** run adds four things (all TTY-only; `--yes`/`--json` surfaces are unchanged):

- **Model picker.** With several loaded models and no `--model`, the wizard's numbered picker
  runs — never a silent first-pick. Non-interactive stays documented behaviour: `--model`, else
  the endpoint's first loaded model.
- **Thinking level, before the confirm screen.** A numbered choice — *model default / off / low /
  medium / high* — that feeds the same `--reasoning-effort` machinery the flag does and is recorded
  identically. `off` sends the upstream-documented `reasoning_effort: none`; support is
  model-dependent, so an engine may ignore or reject it and the run records what the model
  **actually did** (measured from its output), not what was asked. Non-interactive is unchanged:
  the flag, or the model's default.
- **Pre-flight, before the proceed prompt.** What the check is in plain words, how many questions
  against which model, and an **estimated duration range** measured on this machine. The estimate
  is a range, never a single confident number: its low end is the suite's token floor at the
  fastest rate measured, its high end that floor at the slowest rate, multiplied by the suite's
  thinking cost **only when the probe actually observed reasoning** (configuration alone is not
  evidence — an engine may ignore the control). The probe sends the **selected** model's id and
  **discards its first generation**, so a cold weight load never lands in the measured rate. An
  unavailable estimate is reported as unknown, never invented.
- **Honest progress + safe abort.** Progress lines read `question 12/100 · ~1h 40m left` (rolling
  ETA from measured per-item times); raw item ids appear only with `--verbose` (they always ride
  the `--json` payload). `Ctrl-C anytime — partial runs are never uploaded`: SIGINT aborts
  cleanly (exit 130) and uploads **nothing** — a partial suite never scores, so it is never
  submitted, and a started suite is never time-truncated.

`code-v1` runs the identical path: the runner administers the pooled code prompts and uploads
OUTPUTS + timings, and the held-out tests never leave the server (they are the answer key — see
`executor/README.md`). Because code-v1 is scored asynchronously (queue → ephemeral container →
judge) rather than inline, the runner **polls it to a final state** on a longer per-benchmark
window (15 min vs. core-v1's 5); if the window closes first it says so and tells you to report the
run — it never assumes success. code-v1 uses the same sampler defaults as core-v1 (temperature 0,
fixed seed): no in-repo code-v1 spec prescribes a different profile, and the runner does not invent
one. **Consent** is shown once
(state in `~/.crowdbench`); `--agree-tos` attests a human has seen and agreed to the terms (what
uploads; the public CC-BY-NC-SA 4.0 dataset; immutability) and is required to upload non-interactively.

## Commands

| command | what it does |
|---|---|
| *(no arguments, on a TTY)* | the **guided wizard** — zero-file, zero-flag contribution flow |
| `<benchmark-id>` | run one benchmark end-to-end (the flag path); `core-v1` / `core-v1-rc` / `code-v1` / `throughput-std-v1` in this phase |
| `bridge lmeval:<task>:<variant>` | run/parse an allowlisted lm-eval suite → validate → upload |
| `detect` | scan common local serving ports; report apps + loaded models (by route signature) |
| `inventory` | list local models (Ollama / LM Studio / HF cache / user dirs) + benchmark tooling |
| `doctor` | probe + connectivity + state report for support |
| `cache list\|clear` | manage the content-addressed dataset cache |

Key flags (dual-surface parity — every wizard prompt has a flag): `--endpoint`, `--api-base`,
`--dataset` (omit to auto-fetch), `--full` (answer the complete pool instead of the server-issued
subset), `--model`, `--model-path`, `--reps N`, `--seed`,
`--reasoning-effort none|minimal|low|medium|high`, `--reasoning-budget`, `--agree-tos`, `--minutes`,
`--hf-repo`/`--hf-revision`/`--unattributed`, `--endpoint-api-key`
(env `CROWDBENCH_ENDPOINT_API_KEY`; never logged/uploaded/stored), `--dry-run`,
`--no-upload`/`--save`/`--upload-only`, `--json`, `--yes`, `--verbose` (raw item ids on the
progress lines), `--sudo-probes`, `--ports`. Auth uses `CROWDBENCH_API_KEY` (else an anonymous
submitter is minted and cached).

## Trust posture

- **Outputs only.** The runner uploads model outputs + telemetry; it never sees, computes, or
  transmits answer keys. Benchmark item content is inert data — never a tool, never agentic.
- **Dataset integrity.** A dataset's sha256 must match the benchmark's `runner_spec` pin before
  every run (capture requirement 11); a mismatch refuses to run (a tampered set never runs).
- **Vendored artifacts, byte-matched.** `crowdbench_run/_artifacts/*.json` are copied verbatim
  from `packages/shared/artifacts` so both languages read one source; CI fails on drift
  (`scripts/check_artifact_drift.py` — the third leg of the cross-language drift check).
- **Endpoint key privacy.** `--endpoint-api-key` is forwarded as the target endpoint's
  Authorization header only, never logged/uploaded/stored.

Supply-chain audit trail: `packages/runner/BUILD-TIME-CHECKS.md`.

## Publishing (PyPI Trusted Publishing)

Releases go out through `.github/workflows/publish.yml` using **PyPI Trusted Publishing
(OIDC)** — there is **no long-lived PyPI API token** in this repo, ever. GitHub mints a
short-lived OIDC token per run and `pypa/gh-action-pypi-publish` (pinned by full commit SHA)
exchanges it with PyPI. Three human gates stand between a dispatch and a live release:
`workflow_dispatch` only (no push/tag/schedule), a typed `confirm` input that must equal
`publish-to-pypi`, and `environment: pypi` (owner approval before any publish step runs). A
preceding job builds the sdist+wheel, runs `twine check --strict`, install-smokes both
artifacts, and refuses the placeholder `0.0.x` version line.

**Owner one-time setup — verify on pypi.org.** Create the Trusted Publisher for this project
under *Manage project → Publishing* (or *Your projects → Publishing* for a pending publisher).
These **three fields must match** the workflow exactly, or PyPI rejects the OIDC exchange:

| PyPI Trusted-Publisher field | Must be |
|---|---|
| PyPI project (package) name | `crowdbench-run` |
| GitHub repository + workflow filename | `crowdbench-ai/crowdbench-dev`, workflow `publish.yml` |
| Environment name | `pypi` |

(The owner also configures the `pypi` GitHub Environment's protection rule — required
reviewers — so gate #3 actually pauses for approval.) This PR only prepares the workflow and
docs; wiring the Trusted Publisher and the environment protection rule is an owner action on
pypi.org / GitHub, and nothing publishes during the dark period.

## Development

```
cd packages/runner
python -m pip install -e ".[dev]"
python -m pytest -q
```

The test suite runs on ubuntu/macos/windows in CI (a release gate — cross-platform posture).

---

CrowdBench is a service of Hooch Labs LLC. © 2026 Hooch Labs LLC · Apache-2.0.
