Metadata-Version: 2.5
Name: zhongyitang
Version: 0.1.17
Summary: 忠义堂 — one-command local LLM forge: mirror setup, halogen/Strata install, ModelScope model fetch, multi-user quota gateway.
Project-URL: Homepage, https://github.com/CodeOfMe/zhongyitang
Project-URL: Repository, https://github.com/CodeOfMe/zhongyitang
Project-URL: Issues, https://github.com/CodeOfMe/zhongyitang/issues
Project-URL: Changelog, https://github.com/CodeOfMe/zhongyitang/blob/main/CHANGELOG.md
Author: CodeOfMe
License: MIT
License-File: LICENSE
Keywords: gateway,halogen,inference,llm,local-llm,mirror,modelscope,openai-compatible,quota,rate-limit,strata,ubuntu
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: System Administrators
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: System :: Systems Administration
Requires-Python: >=3.9
Provides-Extra: dev
Requires-Dist: build>=1.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: twine>=5.0; extra == 'dev'
Provides-Extra: models
Requires-Dist: modelscope>=1.9; extra == 'models'
Description-Content-Type: text/markdown

# Zhongyitang (忠义堂)

**English** | [简体中文](README_CN.md)

> *A big scale to divide the silver, a big bowl to share the wine.*

Turn a **fresh Ubuntu box** into a **multi-user local LLM service with quotas**, in one command.

There is only so much compute in the office and everyone wants a piece of it. Zhongyitang is
how you set the rules: everyone can use the models on the box, but each person gets their own
key, their own limits and their own ledger — go over and come back tomorrow, and nobody can
sit on the GPU alone.

```bash
pip install zhongyitang
zyt forge            # mirrors → deps → clone → download → deploy → test → gateway
```

---

## What it actually solves

A freshly installed Ubuntu is several chores away from serving a local model: apt hangs on
archive.ubuntu.com, git clones from GitHub time out, the weights are tens or hundreds of
gigabytes and have to be picked out of ModelScope, the engine comes up with no auth at all,
letting a colleague use it means worrying they will hog the VRAM, and at the end of the month
nobody can say who used how much.

Zhongyitang turns that chain into one tool:

| Stage | Command | What it does |
|---|---|---|
| 1 Mirrors | `zyt mirror set ustc` | Point apt and pip at a domestic mirror (USTC/TUNA/…), **reversibly** (originals backed up, `restore` puts them back) |
| 2 Deps | `zyt bootstrap` | git, python3-venv, build-essential, cmake, podman/bubblewrap… |
| 3 Clone | `zyt repo clone` | halogen (engine + deploy kit) and Strata (consumer-GPU engine), shallow, proxy-capable |
| 4 Weights | `zyt model list` / `download` | Pick from the catalog, download from ModelScope **direct**, sha256-verified, resumable |
| 5 Raise the tent | `zyt deploy up <id>` | Build a launch plan (preview the argv with `plan`), start the engine, wait for ready |
| 6 Test the blade | `zyt test` | `/health`, `/v1/models`, a real inference, first-token latency and tok/s |
| 7 Split the take | `zyt serve` | Multi-user gateway: a key per person, rate and total quotas, per-person token detail |
| 8 Go public | `zyt tunnel named` | A Cloudflare Tunnel and a fixed HTTPS address, **outbound only, no ports opened** |
| 9 Keep the shop | `zyt console start` | Start the **management console**: models, gateway and tunnel together, the rest is clicking |

Steps one through eight can all be done from the shell, but **the one you use every day is
step nine**: open the page, see who is running, who used how much, which model wants an
upgrade, and how to roll back when an upgrade goes bad. The tool itself is something you
download; the page answers four questions — **is it there, is it whole, should it be
upgraded, and how do I go back**.

---

## Install

```bash
pip install zhongyitang          # or
uv tool install zhongyitang
```

To download with the official ModelScope CLI (recommended — this tool prefers it when present):

```bash
pip install "zhongyitang[models]"
# or point at an existing executable
export ZYT_MODELSCOPE=/path/to/envs/dev/bin/modelscope
```

**Zero runtime dependencies** — the standard library only. It cannot refuse to start because
of a version conflict in some package.

---

## Ten-minute walkthrough

### 0. Look at the machine first

```bash
zyt doctor
```

It reports whether the OS is recognised, whether RAM/VRAM is enough, which commands are
missing, whether the GPU driver is complete and whether ModelScope is reachable, and then
recommends a **tier** that fits your memory.

### 1. Mirrors + dependencies

```bash
zyt mirror set ustc      # also tuna / aliyun / huawei / tencent
zyt bootstrap
```

Switching mirrors is reversible: the original `sources.list` / `ubuntu.sources` is moved
whole into a backup directory and `zyt mirror restore` puts it back verbatim. We **never edit
someone else's file** — we add one of our own.

Look before you leap:

```bash
zyt mirror set ustc --dry-run
```

### 2. Clone the upstreams

```bash
zyt repo clone --engine halogen     # or strata / all
```

GitHub direct from China often times out mid-handshake; configure a proxy first:

```bash
zyt config set repo.proxy http://127.0.0.1:7890
```

### 3. Pick and download models

```bash
zyt model list                      # 17 entries with size, memory need and whether it is installed
zyt model list --engine strata
zyt model download strata-iq2-xs    # resumable; verifies size when it finishes
zyt model verify strata-iq2-xs --deep   # add --deep to check sha256 (slow, tens of GB)
```

The repo names and file sizes in the catalog are **verified against ModelScope's official
API**, not written from memory. Downloads are **forced direct, with every proxy variable
stripped** — ModelScope is in China, and going through a proxy only makes it slower or ends
in a TLS timeout:

```bash
zyt model download halogen-27b --api   # built-in downloader instead of the CLI (with sha256)
zyt model files halogen-27b            # see what the repo really holds before choosing
```

**A neat thing about Strata**: four quantisation tiers share a 26.8 GB shard 2 (a lookup
table). Installing a second tier therefore costs only the size of shard 1, not another
60 GB. The tool accounts for that in its disk estimate.

### 4. Deploy and test

```bash
zyt deploy plan halogen-27b     # print the plan only: argv, env, mounts, ports; start nothing
zyt deploy up halogen-27b       # start, wait for ready, run the smoke test
zyt test --stream               # separately: first-token latency + tok/s
zyt deploy status               # who is running, on which port
zyt deploy down
```

`plan` genuinely pays for itself: halogen under a container runtime and under bubblewrap take
completely different arguments, and looking first is cheaper than reading logs afterwards.

### 5. Split the take: users, quotas, billing

```bash
zyt serve                                   # start the gateway, 0.0.0.0:8080 by default
```

In another terminal:

```bash
zyt admin key                               # get the admin key (stored at <state>/admin.key)
zyt user add alice --rpm 60 --concurrency 2 --daily-tokens 2000000
zyt user add bob   --daily-tokens 200000
```

`user add` prints the API key **once**; only its sha256 is kept. Lost it? Issue a new one:

```bash
zyt user key alice --label phone
```

Point a client at it and go:

```bash
curl http://<host-ip>:8080/v1/chat/completions \
  -H "Authorization: Bearer zyt-......" \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"hello"}]}'
```

### 6. Read the books

```bash
zyt admin report --days 14
zyt admin report --user alice --days 30
zyt admin report --csv > usage.csv
zyt admin report --days 30 --export report.json
```

The CLI prints Chinese; the numbers are the point:

```
== 按用户
用户        今日        7 天        14 天       累计         请求数  Key  今日配额
alice       12.4K       88.1K       176.0K      1.2M         412    2    24.6K/2.0M
bob         3.1K        19.7K       41.2K       41.2K        96     1    3.1K/200K
```

The web view is friendlier, and there are **two doors** — don't take the wrong one:

| Door | For | What you do inside |
|---|---|---|
| `http://<host-ip>:8080/` | a normal user with a key | Paste your API key, see **your own** usage, limits and 14-day curve, plus the `base_url` and model name to put in your client |
| `http://<host-ip>:8080/console` | the admin | Paste the admin key, see global KPIs, per-user and per-key detail, daily usage bars, start/stop models, change limits, upgrade assets |

`zyt console open` opens the admin page directly. Keys stay in your own browser's
localStorage. The front page **does not list** ports, model names or endpoint tables — it is
public, so those wait until after login. See “The console” below.

### 7. Go public: Cloudflare Tunnel

Want people outside the LAN to use it? No public IP, no port forwarding, no certificates of
your own:

```bash
zyt tunnel status                                  # see where things stand: installed, logged in, which hostname
zyt tunnel quick                                   # temporary: a random *.trycloudflare.com address, changes on restart
zyt tunnel named --hostname llm.example.com        # permanent: a fixed address on your own domain
zyt tunnel down                                     # stop it (only the process we started)
```

Or do it in one shot from `zyt serve`:

```bash
zyt serve --tunnel named --tunnel-hostname llm.example.com
```

**quick or named?** quick is anonymous, needs no login, and **changes address every time it
restarts** — good for “let me show you for ten minutes”. named binds to your Cloudflare
account and a fixed hostname and survives restarts — **this is the one a deployment should
use**.

The idea is that `cloudflared` opens a tunnel **outbound** from this machine and Cloudflare
receives the public traffic for you:

```
   public users ──HTTPS──▶ Cloudflare ──tunnel──┐
                                                 ▼
                                     cloudflared (local, outbound only)
                                                 │
                                                 ▼
                                      gateway :8080 (auth + quotas)
                                                 │
                                                 ▼
                          engine :8731 / :8730 (loopback, no auth)
```

A few deliberate choices:

- **Publish the gateway port only.** `:8730` (the engine's own protocol) and `:8731` (the
  upstream front end) have **no authentication at all**; they never enter ingress — that is
  written into the code.
- **Never touch someone else's config.** Before writing `~/.cloudflared/config.yml` we check
  whether it is ours; if not we refuse to overwrite unless `--force`, and always leave a
  `.bak-<timestamp>`. Same for the systemd unit.
- **Stop only what we started.** `tunnel down` trusts only its own pidfile and never scans by
  name or port — there may be a hand-made tunnel on this box.
- **Do not cut long streams.** The generated config carries `keepAliveTimeout: 90s`, because
  the first token of a `stream: true` request can take minutes.
- **`--create` is the only thing that touches the cloud.** `status` / `quick` / `named`
  (without `--create`) are entirely local and read-only.

**One thing to be clear about: a tunnel is reachability, not authentication.** Whoever learns
your hostname is standing at your gateway's door, with only the Bearer key and the quotas in
between. If the audience is bigger than a few trusted people, add
[Cloudflare Access](https://developers.cloudflare.com/cloudflare-one/policies/access/) on top.
`zyt tunnel named` prints exactly that reminder.

---

## The watchdog: the gateway, the model and the tunnel will not come back on their own

The gateway that `zyt console start` launches is **detached from the terminal**, and the
engine is a bwrap process group — neither is a systemd unit, so nothing brings either back,
and the public endpoint stays 502 until a human notices. The only link in the chain with
`Restart=always` is cloudflared itself. The watchdog is the missing piece.

```bash
zyt watchdog status        # last pass + unit state
zyt watchdog once          # run one pass by hand (--dry-run decides but changes nothing)
zyt watchdog run           # stay in the foreground (this is what the systemd unit runs)
zyt watchdog install       # write and enable the systemd --user unit (zyt-watchdog.service)
zyt watchdog uninstall
```

Why it is as conservative as it is:

- A model that is **cold-loading** (minutes) is not a failure: it is restarted only when the
  process is truly gone, or when the health check has failed for 20 minutes straight.
- With a download in flight, an UP-but-hung gateway waits 30 minutes before a hard restart —
  restarting would orphan the download. A gateway that is simply **gone** is started
  immediately, because the download's own helper process is still writing.
- Every decision goes into `logs/watchdog.log`, including the reasons it decided *not* to act.

**Download resume**: orphaned `.incomplete` downloads left behind by a gateway restart are
adopted; if one stops growing for three minutes and nobody holds it open (we scan `/proc` for
handles) it runs `zyt assets update` to continue, with a five-minute backoff that doubles up
to one hour and **never gives up**; if another process is still writing, we do not touch it —
two writers on one file is the thing this tool fears most.

**The tunnel check is functional**: a green `systemctl is-active` does not mean reachable
(this project's reference box was bitten by a fake-IP proxy: the unit was healthy while every
TLS handshake to the Cloudflare edge was reset and the public site returned 502).
The watchdog probes the public `/health` every 15 seconds; two minutes of failure while the
gateway itself is healthy restarts cloudflared. When the gateway is unhealthy too, the tunnel
is left alone — it is not the tunnel's problem then.

`zyt watchdog install` also writes the current `ZYT_GATEWAY_UPSTREAM` / `ZYT_UPSTREAM_KEY`
into the unit: the gateway is detached, so the watchdog needs to know where to point it and
which key to use when it has to start it again.

---

## The console

```bash
zyt console start            # models + gateway + console (--no-models: gateway only, --no-wait: don't wait for ready)
zyt console url              # the address
zyt console key              # the admin key
zyt console open             # open a browser
zyt console logs gateway     # who is noisy (or a specific model)
zyt console status
zyt console stop
```

Open the address, paste the admin key, and there are eight tabs:

| Tab | What it does |
|---|---|
| Overview | Version, hardware, who is running, is the upstream reachable, tunnel state |
| Models | Every model in the directory: installed, running, on which port; start/stop |
| Assets | The workspace list: is it there, is it whole, should it be upgraded; verify / install / update / rollback |
| Users | Per person: today / 7 days / 14 days / lifetime tokens, request count, key count, quota used |
| Keys | Per-key detail; rpm, tpm, daily tokens, total tokens, concurrency and expiry are all editable |
| Ledger | Daily usage bars, recent request stream |
| Settings | Workspace, model, repo, Strata data directories and proxy |
| Access | Gateway address, tunnel address, what to put in a client |

Narrow screens first: single column on a phone, thumb-sized buttons, tabs that scroll
sideways instead of folding into a knot — you can glance at whether the model is up and who is
eating the GPU from the sofa. The page stays **self-contained** and pulls no external
resources: the machine it runs on may sit on an isolated network, and a page that needs the
internet to render would break exactly when you need it.

**Eleven languages**: Chinese / English / Japanese / Korean / French / German / Dutch
(including Belgian) / Spanish / Portuguese / Italian / Russian. The first visit picks the
language from the browser (anything unknown falls back to English); there is a manual switch
in the top right, and the choice is stored on the account, so it follows you to another
machine. Dynamic text, confirmation dialogs, code sample blocks and the durations inside task
rows are translated too.

### Assets: one list, two sources, four actions

Everything a deployment needs lives under **one workspace directory**, is described by
**one list**, and is driven by **four actions**:

```bash
zyt assets list                    # the list and its state (--check also asks the repos how far behind)
zyt assets verify                  # is it there, is it whole (--deep to actually check sha256)
zyt assets install                 # fill in what is missing
zyt assets update                  # upgrade to the latest
zyt assets rollback model:halogen-27b
zyt assets jobs                    # progress and logs of background jobs
```

The two sources are kept apart because weights and binaries **are not the same kind of thing**:

| Source | What “version” means | Snapshot | Rollback |
|---|---|---|---|
| `model` weights (ModelScope) | the pile of files on disk | a `{path: size}` listing | delete what this run added |
| `repo` executables (git clone) | a commit | a sha | `git reset --hard` |

Downloads take hours, so **no action runs inside an HTTP request**: you click, get a job number
immediately, a background thread works and appends to a size-capped log, and the page polls it.
A failing action **restores the snapshot before reporting the failure**, so “it broke” does not
leave a half-updated workspace behind.

`verify` is local-only by default (existence + size), because hashing 200 GB is slow; add
`--deep` when it matters. The verdict only counts **weight files**: `.part` does not count, and
the other 140 GiB of shards sharing a Strata directory are not yours.

### Running several models at once

`halogen.sh` solves “one model per machine”. Zhongyitang adds a dimension: run several at
once, know who is on which port, and let the gateway route each request by its `model` name.
There is exactly one source of truth — `<state>/runtime/models.json` — read by the console, the
gateway and `zyt model ps`; nobody guesses from `ps`.

The port allocator always deals **two numbers per instance** (`:8730` the engine's own
protocol, `:8731` the OpenAI front end). In container mode the second number goes unused, but
**one piece of logic beats two** — bwrap mode shares the host network namespace, where both
are genuinely needed. Every instance leads its own process group, so teardown is one `killpg`
that takes the whole tree, because Strata forks a `serve/server.py` and killing only the parent
leaves an orphan holding the VRAM.

---

## How quotas work

Four dimensions, **every key can override its user's defaults, and 0 means unlimited**:

| Dimension | Field | Effect |
|---|---|---|
| Daily tokens | `daily_tokens` | Resets on the **local calendar day**, not UTC, so nobody farms a timezone |
| Total tokens | `total_tokens` | Lifetime allowance; good for temporary accounts |
| Requests/min | `rpm` | Stops retry loops spinning like a flywheel |
| Tokens/min | `tpm` | Stops one 250k-token monster prompt |
| Concurrency | `concurrency` | **The important one on a single-slot engine**: two 200k-token requests at once do not each run half as fast, they strangle each other |

Token counts come from the engine's returned `usage` — **client-reported numbers are not
trusted**. When the engine gives no usage, the gateway estimates from text length and records
that honestly. Every request lands in SQLite, and the admin report and the quota verdict read
the same data, so “the report says I'm fine, the gateway says I'm over” cannot happen.

You do not have to memorise command-line flags for these — the console's **Keys** tab edits
them per key, and the next request uses the new limits. `zyt user limits <name>` is the
command-line equivalent (note: with no arguments it **only shows the current state and writes
nothing**; pass values to change, pass `0` to lift a limit).

The admin master key (`ZYT_ADMIN_KEY` or `<state>/admin.key`) bypasses every limit and can read
`/admin/api/*`.

---

## Command cheat sheet

```
zyt forge                        the whole pipeline in one go
zyt doctor [--quick]             environment check + tier recommendation
zyt mirror show|set|restore      package mirrors (--dry-run to preview)
zyt bootstrap                    install git / build / container dependencies
zyt repo list|clone|update|status|rollback   clone or update halogen and Strata
zyt assets list|status|verify|install|update|rollback|jobs   workspace assets
zyt model list|search|files|download|verify|installed|prune
zyt model run|stop|restart|ps        start/stop one model (several at once: use the console)
zyt deploy plan|up|down|status|configs
zyt test [--stream]              smoke test and throughput
zyt user add|list|show|limits|rm|key|keys|rm-key
zyt admin report|key|prune|db    reports, master key, cleanup, ledger location
zyt console start|stop|restart|status|url|key|logs|open      models + gateway + console
zyt serve                        the multi-user gateway alone
zyt tunnel status|quick|named|down   Cloudflare Tunnel: publish the gateway on a fixed HTTPS address
zyt watchdog status|once|run|install|uninstall   keep gateway/model/tunnel up, resume downloads
zyt config show|get|set|path     read and write configuration
```

Except for `zyt forge` (the one interactive command that reports progress as it goes), every
subcommand supports `--json` so it can go into a script. Exit codes: `0` fine, `1` business
failure, `2` usage error. Note that `doctor`'s `1` is a **verdict** (is this machine good
enough), not a failed command.

---

## Configuration

Configuration is JSON, merged in three layers: **defaults ← `$ZYT_HOME/config.json` ←
environment variables**.

```bash
zyt config show
zyt config set gateway.port 8081
zyt config set engine.ctx 131072
zyt config set assets.dir /data/llm          # workspace: model files + engine source trees
zyt config set model.dir /data/llm/models
zyt config set repo.dir /data/llm/repos
zyt config set engine.strata_data_dir /data/llm/strata
zyt config path
```

The **workspace directory** (`assets.dir`) is wherever you decided to keep things; the other
three (`model.dir`, `repo.dir`, `engine.strata_data_dir`) fall back under the state directory
when left empty. If a layout already exists somewhere else — models under
`/srv/llm/models`, Strata data under `/srv/llm/strata`, engine source trees under
`/srv/llm/src` — point all four at it and change nothing else:

```bash
zyt config set assets.dir /srv/llm
zyt config set model.dir /srv/llm/models
zyt config set repo.dir /srv/llm/src
zyt config set engine.strata_data_dir /srv/llm/strata
```

### No proxy by default

A desktop proxy client (Clash/mihomo and friends) exports `HTTP_PROXY` / `all_proxy` into the
environment, and every child process inherits them. But zhongyitang's downloads, clones and
health probes should not take that detour: ModelScope is faster direct, and a loopback probe
through a proxy times out for certain. So **zhongyitang strips the proxy variables from itself
and every child by default**; set `ZYT_USE_PROXY=1` to go back to “follow the environment”. If
you really do need a proxy, configure it per component:

```bash
zyt config set repo.proxy http://127.0.0.1:7897     # cloning upstream repos
zyt config set model.proxy http://127.0.0.1:7897    # ModelScope (usually best left empty/direct)
zyt config set assets.proxy http://127.0.0.1:7897   # asset downloads/updates
```

| Environment variable | Effect |
|---|---|
| `ZYT_HOME` | State directory (default `~/.local/share/zhongyitang`) |
| `ZYT_ADMIN_KEY` | Override the admin master key |
| `ZYT_MODELSCOPE` | Point at a modelscope executable |
| `ZYT_NO_SYSTEM=1` | **Dry-run only**: write no /etc, run no apt (essential in CI/containers) |
| `ZYT_MIRROR_APT` / `ZYT_MIRROR_PIP` | Override the mirror |
| `ZYT_ASSETS_DIR` | Workspace directory (common root of model files and engine source trees) |
| `ZYT_ASSETS_PROXY` / `ZYT_REPO_PROXY` | Proxy for fetching upstream repos (unset = direct) |
| `ZYT_MODEL_DIR` | Model weights directory |
| `ZYT_REPO_DIR` | Engine source tree directory |
| `ZYT_STRATA_DATA_DIR` | Strata data directory |
| `ZYT_GATEWAY_BIND` / `_PORT` / `_UPSTREAM` | Gateway bind and upstream |
| `ZYT_UPSTREAM_KEY` | Bearer sent to the upstream engine (used when proxying) |
| `ZYT_CONSOLE_BIND` / `ZYT_CONSOLE_PORT` | Console bind (empty = follow the gateway) |
| `ZYT_CLOUDFLARED` | Point at a cloudflared executable (default: PATH, then `~/.local/bin`, …) |
| `ZYT_CLOUDFLARED_CONFIG` | Override the cloudflared config path (default `~/.cloudflared/config.yml`) |
| `ZYT_UNIT_DIR` | Override the systemd user unit directory (default `~/.config/systemd/user`) |

Directory layout:

```
~/.local/share/zhongyitang/
├── config.json         configuration
├── zhongyitang.db      ledger (SQLite: users / keys / usage)
├── admin.key           admin master key (0600)
├── tunnel.json         tunnel state (mode, address, pid)
├── cloudflared-quick.yml  quick-tunnel config (catch-all)
├── cloudflared.pid     the tunnel process we started
├── models/             model weights (or wherever `model.dir` points)
├── repos/              cloned halogen and Strata (or wherever `repo.dir` points)
├── runtime/            runtime state
│   ├── models.json     control plane: who is running, on which port, pid
│   ├── assets-jobs.json  asset job (verify/install/update/rollback) progress and logs
│   ├── deploy/         deployment records (one per model: argv / env / mounts / ports)
│   ├── files/          pid and operation locks
│   └── gateway.log     gateway log
├── cache/              download staging
└── logs/
```

---

## Supported models

Seventeen entries in three groups, all verified against ModelScope.

**halogen (Strix Halo engine, single-file `.hgn` weights)**

| id | Size | Note |
|---|---|---|
| `halogen-flash-v2` | 62.1 G | Qwen3.8-Flash-Next v2 checkpoint |
| `halogen-flash-ngram` / `-ht43` | 47.7 / 53.7 G | variants of the same family |
| `halogen-flash-w4b` | 115.6 G | highest precision, also the largest |
| `halogen-27b` | 33.4 G | Qwen3.8-27B; 32K prompt end to end in 66 s |
| `halogen-flash-mtp` | 1.4 G | speculative-decoding draft layer, speeds up any checkpoint |
| `halogen-flash-vision` | 0.8 G | vision encoder; only needed for image input |

**Strata (consumer GPUs, GGUF shards)**

| id | Size | Memory | Note |
|---|---|---|---|
| `strata-iq2-xs` | 63.4 G | 42 G | **recommended**, the balance of speed and quality |
| `strata-q2-0` | 61.9 G | 40 G | fastest |
| `strata-iq3-xxs` | 70.6 G | 50 G | better quality |
| `strata-iq3-s` | 77.9 G | 58 G | public evals match the base model |
| `strata-coder` | 29.6 G | 32 G | code-focused; noticeably worse in Chinese |
| `strata-swift-iq2-xs` | 63.4 G | 42 G | third-party fine-tune, answers faster |
| `strata-unsloth-ud-q4-k-xl` | 103.4 G | 48 G | experimental 4-bit, slow |

**NPU sidecars**: `npu-embedding` / `npu-reranker` / `npu-decider`, for `/v1/embeddings`.

---

## Gateway endpoints

| Endpoint | Auth | Note |
|---|---|---|
| `POST /v1/chat/completions` | user key | chat, supports `stream: true` |
| `POST /v1/embeddings` | user key | embeddings (needs an NPU sidecar) |
| `GET /v1/models` | user key | model list |
| `GET /usage` | user key | your own usage and limits |
| `GET /health` `GET /ready` | none | liveness and upstream readiness |
| `GET /` | none | **user portal**: the login page is public; paste a key to see your own usage |
| `GET /admin` `GET /console` | admin key | management console (eight tabs) |
| `GET /admin/api/{env,models,access,settings,assets,jobs,overview,series,recent,user,users,keys}` | admin key | JSON the console reads |
| `POST /admin/api/models/<id>/{start,stop}` | admin key | start/stop a model (`DELETE` is a synonym for stop) |
| `POST /admin/api/assets/<id>/{verify,install,update,rollback}` | admin key | submit an asset job, **a job number comes back immediately**; poll `GET /admin/api/jobs?id=<job>` |
| `POST /admin/api/settings` | admin key | change workspace/model/repo/Strata directories and proxy |
| `POST /admin/api/{users,users/<id>,keys,keys/<id>}` | admin key | create users, issue keys, change limits |
| `POST /admin/api/keys/<id>/{enable,disable}` · `DELETE /admin/api/{users/<id>,keys/<id>}` | admin key | disable / re-enable / delete (the last admin cannot be deleted) |

Over quota returns `429` with `{"error":{"type":"insufficient_quota"}}`; over rate returns `429`
with `rate_limit_exceeded` and a `Retry-After`.

---

## About security

- API keys are stored as sha256 only; after a leak they **cannot be recovered, only reissued**.
- Comparisons use `hmac.compare_digest`, and keys are always masked in logs.
- The gateway binds `0.0.0.0` by default; the upstream engine ports (8731/8730) **always bind
  loopback**. Do not expose an unauthenticated engine directly. To go public use
  `zyt tunnel named` — it publishes the gateway port only, the engine ports never enter
  ingress, and the tunnel is **outbound only**, opening no listening port on the internet.
- A tunnel is not authentication: whoever learns the hostname can reach your gateway's door,
  stopped only by the Bearer key and the quotas. A wide audience deserves Cloudflare Access.
- The pages are a public HTML shell with **no data inside** — the key is entered on the page, so
  the page must load first. Both pages obey that, and the front page (`/`) **does not even
  describe the server**: no ports, no endpoint table, no model names, and it will not print the
  address you reached it at (LAN IP included). That only appears after a key is pasted and
  `/usage` answers.
- All data and management actions (`/admin/api/*`) require the admin key. A pasted key stays in
  **your own browser's localStorage**; no cookies, no third parties; requests always carry an
  `Authorization` header and a failure drops back to the lock screen.
- A client that connects and disconnects **no longer writes a stack trace to the log**.
  Otherwise anyone (authenticated or not) could fill the log by connecting and dropping in a
  loop, burying the lines that matter.
- `admin.key` is mode 600.
- Neither the engine nor the gateway runs as root. Only mirrors and package installation need
  root.

---

## Development

```bash
git clone https://github.com/CodeOfMe/zhongyitang
cd zhongyitang
python -m venv .venv && . .venv/bin/activate
pip install -e ".[dev]"
PYTHONPATH=src ZYT_NO_SYSTEM=1 pytest -q
```

The tests touch no network, no `/etc` and no real model: the gateway suite **starts a real
gateway** against a fake upstream and asserts authentication, quotas, rate limiting, SSE
pass-through and accounting over HTTP.

```bash
python -m build          # sdist + wheel
twine check dist/*
```

**One extra step inside this project's WorkBuddy environment.** That environment injects a
safe-delete shim into Python (`.../cli/vendor/shim` on `PYTHONPATH`) which blocks any `rmtree`
of more than 50 entries and calls `SystemExit(1)`. pytest clears its basetemp every run and
trips it as soon as the tree is big, so **every fixture that uses a temp directory blows up**
and 364 tests report 364 ERRORs — it looks like the code is broken, but not one line ran. Run
it like this:

```bash
env -u PYTHONPATH CODEBUDDY_SAFE_DELETE_ENABLED=0 pytest -q --basetemp=~/.cache/zyt-pytest
```

---

## License

MIT.
