Metadata-Version: 2.4
Name: gpu-broker
Version: 0.5.0
Summary: Share one GPU between LLM servers and ComfyUI: a queue that swaps models in and out on demand.
Author: Ashwin Philips
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/emergenthq-net/gpu-broker
Project-URL: Repository, https://github.com/emergenthq-net/gpu-broker
Project-URL: Issues, https://github.com/emergenthq-net/gpu-broker/issues
Project-URL: Releases, https://github.com/emergenthq-net/gpu-broker/releases
Project-URL: Changelog, https://github.com/emergenthq-net/gpu-broker/blob/main/CHANGELOG.md
Project-URL: Documentation, https://github.com/emergenthq-net/gpu-broker/tree/main/docs
Keywords: gpu,llm,comfyui,inference,scheduler,mcp
Classifier: Programming Language :: Python :: 3
Classifier: Framework :: FastAPI
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Operating System :: POSIX :: Linux
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: fastapi<1,>=0.115
Requires-Dist: uvicorn<1,>=0.30
Requires-Dist: pyyaml<7,>=6.0
Provides-Extra: download
Requires-Dist: huggingface_hub>=0.34; extra == "download"
Provides-Extra: mcp
Requires-Dist: mcp<3,>=2.3; extra == "mcp"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: httpx>=0.27; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: mypy>=1.11; extra == "dev"
Requires-Dist: types-PyYAML; extra == "dev"
Requires-Dist: openai>=1.60; extra == "dev"
Requires-Dist: anthropic>=0.50; extra == "dev"
Requires-Dist: mcp<3,>=2.3; extra == "dev"
Dynamic: license-file

# gpu-broker

<p align="center">
  <picture>
    <source media="(prefers-color-scheme: dark)" srcset="docs/img/logo-dark.png">
    <img src="docs/img/logo-light.png" alt="gpu-broker" width="320">
  </picture>
</p>

## TLDR: just install it

```bash
pip install gpu-broker      # or: uvx gpu-broker setup
gpu-broker setup
```

It finds your GPU and model servers, writes the config, starts the service and opens the dashboard.

---

gpu-broker lets a chat model, an image generator and a video generator take turns on one
graphics card, switching between them for you as requests arrive.

[![PyPI](https://img.shields.io/pypi/v/gpu-broker.svg)](https://pypi.org/project/gpu-broker/)
[![CI](https://github.com/emergenthq-net/gpu-broker/actions/workflows/ci.yml/badge.svg)](https://github.com/emergenthq-net/gpu-broker/actions/workflows/ci.yml)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
![Python 3.12+](https://img.shields.io/badge/python-3.12%2B-blue.svg)

## Who it's for

**You, with one GPU and more work than fits on it.** Your background agents keep a local LLM
busy all day. You chat with it too. Now and then you want an image or a short video. The chat
model and the image model don't fit in the card's memory together.

Without gpu-broker you stop the LLM, start ComfyUI, make the image, and restart the LLM by
hand; your agents stall until you do. With it, you just send the request.

It also fits:

- **A small team sharing one workstation.** Everyone gets one address; nobody switches models by hand.
- **Scripts and agents** that need several kinds of model (chat, image, video, 3D) from one machine.
- **A home lab** with one good GPU.

## What it does

- **One address for everything.** Chat, image and video requests all go to gpu-broker.
- **Takes turns on the card.** For an image or video job it:
  - lets the chats in progress finish,
  - stops the chat model,
  - runs the job,
  - queues whatever arrives meanwhile.
- **Puts the chat model back by itself** once the card has been quiet for a couple of minutes.
- **Keeps people ahead of batch work.** Your own chats skip the queue and stream token by token.
- **Shows what is happening.** A web dashboard shows what is loaded, what runs, what waits, and every switch.
- **Speaks the OpenAI (Chat Completions and Responses) and Anthropic APIs,** so Open WebUI, the official SDKs and the Codex CLI connect unchanged.
- **Cloud first, local when it fails (optional).** With `upstreams:` routes, `claude-*` and `gpt-*` go to the real provider with the client's own key, and fall back to a local model on an outage, a hung connection or used-up quota; a circuit breaker sends them back when the provider recovers ([docs/failover.md](docs/failover.md)).

## See it

![The dashboard: the chat model is loaded and answering requests](docs/img/overview.png)
*The chat model is loaded and answering. The live charts show the card's memory, load, power
and temperature.*

![A video job running after the chat model was stopped, with chat requests waiting in the queue](docs/img/video-queue.png)
*A video was requested. The chat model was stopped to make room; chats that arrive wait in the
queue until the video is done.*

![The model list, with buttons to borrow the GPU for an image or video model](docs/img/models.png)
*Every model it can run. You can borrow the whole GPU for hands-on work in ComfyUI; it goes
back to the chat model when you are done.*

![The event log: the chat model stopped for an image, then restored once the card went quiet](docs/img/events.png)
*The event log, newest first: the chat model stopped for an image, then came back on its own.*

## Try it in 30 seconds

No GPU needed. The demo runs the real dashboard and API on a simulated card, with simulated
users. You need Python 3.12 or newer.

```bash
pip install gpu-broker
gpu-broker demo
```

Or without installing: `uvx gpu-broker demo`.

- The dashboard opens in your browser. Over SSH, or with `--no-browser`, open the printed link.
- Within two minutes, someone asks for a video and the chat model steps aside.
- `gpu-broker demo --quiet` leaves out the simulated users, so you can send your own requests.
  It prints a `curl` line to start from.
- Nothing real runs: no model is downloaded, and the "images" are placeholders.

## Quickstart

```bash
pip install gpu-broker               # add [download] to fetch models: pip install 'gpu-broker[download]'
gpu-broker setup                     # or first see what it would do: gpu-broker setup --dry-run
```

`setup` asks no questions. It:

- **finds the GPU** (NVIDIA or AMD) and its memory;
- **finds your model servers**: llama.cpp, vLLM, Ollama and ComfyUI on their usual ports, and the
  systemd units or Docker containers that run them;
- **writes** `config.yaml` and `catalog.yaml` for what it found (the starter files if it found
  nothing), and a new API token in `broker.env` (mode 600);
- **installs and starts the `gpu-broker` service**: a user service if your model servers are
  `systemctl --user` units, else a system service when it has root or passwordless sudo. The
  broker never runs as root. Otherwise it prints the exact `serve` command, and runs it for you
  when you're at a terminal;
- **runs `check`**, waits until the broker answers, and **opens the dashboard** (not over SSH).

Running it again is safe: it keeps every file it finds, and the token. Details, flags and what
each estimate means: [docs/setup.md](docs/setup.md).

## Manual setup

For full control, set it up by hand: pick how your model servers run.

| your model servers are | route |
|---|---|
| systemd units on this machine | [systemd](#systemd) |
| Docker containers | [Docker compose](#docker-compose) |
| systemd units in Proxmox LXCs | [Proxmox](#proxmox) |

### systemd

Create the model servers' units as you normally would. Then:

```bash
sudo gpu-broker init                 # writes /etc/gpu-broker/{config,catalog}.yaml, creates its folders
sudoedit /etc/gpu-broker/catalog.yaml   # your units, endpoints and model sizes
gpu-broker check                     # validates both files, prints the driver and the unit allowlist
sudo BROKER_TOKEN=$(openssl rand -hex 24) gpu-broker serve
```

- `init` never overwrites existing files (`--force` does). `--dir` writes elsewhere.
- It listens on `127.0.0.1:8095` (`server.host`, `server.port`).
- The broker runs `systemctl start/stop` on the catalog's units.
  - Run it as root, or set `driver.sudo: true` with a sudoers rule limited to those units.
- To run it as a service: [`examples/systemd/gpu-broker.service`](examples/systemd/gpu-broker.service).
  - Keep its `TimeoutStopSec` above `server.graceful_shutdown_s` (default 10 s).

### Docker compose

This route needs a clone, for the compose files in `examples/docker`. It also needs Docker and
the NVIDIA Container Toolkit.

```bash
git clone https://github.com/emergenthq-net/gpu-broker && cd gpu-broker/examples/docker
mkdir -p conf data models/llama-3.1-8b && cp config.yaml ../catalog.yaml conf/
sudo chown -R 10001:10001 conf data          # the broker runs as uid 10001 and rewrites catalog.yaml
echo "BROKER_TOKEN=$(openssl rand -hex 24)" > .env
echo "DOCKER_GID=$(getent group docker | cut -d: -f3)" >> .env
pip install huggingface_hub                  # provides the `hf` CLI, for the one-off model fetch
hf download bartowski/Meta-Llama-3.1-8B-Instruct-GGUF Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
  --local-dir models/llama-3.1-8b
docker compose up -d --build
```

- It runs the broker, llama.cpp's `llama-server` and ComfyUI.
- The broker starts and stops the other two through the Docker socket. It never creates or removes containers.
- The dashboard is at http://localhost:8095/dash. It asks once for the token from `.env`.
- Image and video models need their files in ComfyUI's model folders (the `comfy-models` volume).
  - The comments in [`examples/catalog.yaml`](examples/catalog.yaml) name the files each entry expects.

### Proxmox

The broker runs in its own LXC or VM. The model servers are systemd units in other containers.

1. Start from [`examples/config.proxmox.yaml`](examples/config.proxmox.yaml).
2. On the host, install [`host/gpu-broker-ctl`](host/gpu-broker-ctl) as the broker key's forced command
   (see [Security](#security)).
3. Install [`host/gpu-broker-gpu`](host/gpu-broker-gpu), its GPU reader, next to it.

## Use it

```bash
# after setup the token is in broker.env (/etc/gpu-broker, or ~/.config/gpu-broker without root)
export BROKER_TOKEN=$(sudo sed -n 's/^BROKER_TOKEN=//p' /etc/gpu-broker/broker.env)
T="Authorization: Bearer $BROKER_TOKEN"
curl -s localhost:8095/v1/chat/completions -H "$T" -H 'Content-Type: application/json' \
  -d '{"model":"llama","messages":[{"role":"user","content":"hi"}]}'
curl -s localhost:8095/v1/jobs -H "$T" -H 'Content-Type: application/json' \
  -d '{"model":"wan2.2-5b","prompt":"a fox in snow","wait":true}'
```

More examples and every route: [docs/api.md](docs/api.md).

## Connect your apps

```bash
gpu-broker connect --url http://gpu-host:8095     # or click Connect apps on the dashboard
```

- **`gpu-broker connect`** finds the tools on your machine (your shell, Continue, Cline, Aider,
  Codex, Open WebUI) and points them at your local model. No settings to edit.
  - `--dry-run` shows the plan first; `gpu-broker clients` lists what was found and connected.
  - `gpu-broker disconnect` puts everything back, from backups taken before the first change.
- **Any other app** built for the OpenAI (Chat Completions or Responses) or Anthropic API works
  by changing only the base URL and key.
- Details: [docs/drop-in.md](docs/drop-in.md).

### Use it from Claude, ChatGPT or Codex (MCP)

- With the `mcp` extra (`pip install 'gpu-broker[mcp]'`), the broker is also an MCP server:
  assistants can generate and edit images, make videos and 3D scenes, and ask the local model.
- Remote clients and ChatGPT connectors use `/mcp` on the broker; apps that launch a command use
  `gpu-broker mcp`. `gpu-broker connect` registers it in Claude Code, Claude Desktop and Codex.
- Details: [docs/mcp.md](docs/mcp.md).

## Documentation

| page | covers |
|---|---|
| [docs/setup.md](docs/setup.md) | `gpu-broker setup`: what it detects, what it writes, flags |
| [docs/catalog.md](docs/catalog.md) | the catalog: models, substitution, input files, bundled ComfyUI templates |
| [docs/exec-recipes.md](docs/exec-recipes.md) | command-line models (`runner: exec`): recipes, timeouts, GPU holds |
| [docs/api.md](docs/api.md) | every route, request fields, headers |
| [docs/drop-in.md](docs/drop-in.md) | using it in place of the OpenAI and Anthropic APIs |
| [docs/failover.md](docs/failover.md) | cloud first, local on failure: routes, the circuit breaker, what fails over |
| [docs/mcp.md](docs/mcp.md) | the MCP server: tools, `/mcp` and `gpu-broker mcp`, registering it |
| [docs/hardware.md](docs/hardware.md) | host drivers, supported GPUs, choosing the card |
| [ARCHITECTURE.md](ARCHITECTURE.md) | module layout and invariants |
| [SECURITY.md](SECURITY.md), [docs/threat-model.md](docs/threat-model.md) | security model and threat model |

### Command-line models

- Some models are a program, not a ComfyUI graph: image-to-3D tools, for example.
- gpu-broker runs them from a **recipe** file the host's administrator writes.
- A request names the recipe but never carries a command.
- Details: [docs/exec-recipes.md](docs/exec-recipes.md). Example: [`examples/recipes/sharp.recipe`](examples/recipes/sharp.recipe).

## How it works

```mermaid
sequenceDiagram
    autonumber
    actor U as You (chat)
    actor A as Your agent
    participant B as gpu-broker
    participant L as LLM server
    participant C as ComfyUI
    U->>B: POST /v1/chat/completions (stream)
    B->>L: resident, slot free: forward directly
    L-->>U: tokens, streamed as generated
    A->>B: POST /v1/jobs {model: video, prompt}
    B->>B: queue the job
    B->>B: close the LLM pool, drain in-flight calls
    B->>L: stop the unit
    B->>C: run the video graph
    C-->>B: outputs
    B-->>A: job done, with output file URLs
    Note over B: queue idle for idle_restore_s
    B->>C: POST /free
    B->>L: start the unit, wait for /health
    U->>B: next chat is served directly again
```

- **One GPU thread, fair order.** Models switch only between jobs, never mid-job.
  - Waiting jobs go interactive first, then by requester: whoever has used the least expected GPU time goes next.
  - Each requester's jobs keep their order, and an idle caller cannot bank credit.
  - A background job that has waited `scheduler.max_wait_s` (default 1 h) counts as interactive.
  - `scheduler.policy: fifo` restores strict arrival order.
- **Every LLM call holds a pool slot.** Before a switch, the pool closes and in-flight calls finish.
- **People first.** Background calls may fill `slots - reserved_interactive` slots; people may use them all.
- **Safe switches.** ComfyUI frees its weights before an LLM starts; the LLM stops before a ComfyUI job.
  - Health is re-checked, not assumed, since someone else may stop things.
- **Idle restore.** After `idle_restore_s` with an empty queue, `defaults.resident` comes back.
- **Log.** Every state change is a SQLite row and a JSONL line.
- **Replay.** `gpu-broker replay /var/log/gpu-broker/events.jsonl` runs your own log through the queue
  rules in simulated time and prints the wait per requester and the residency switches, next to what
  the broker really did. It contacts nothing; it is how scheduler changes are measured.

### Compared with llama-swap

[llama-swap](https://github.com/mostlygeek/llama-swap) is excellent if all you run is
OpenAI-compatible LLM servers. gpu-broker covers what it doesn't:

| | llama-swap | gpu-broker |
|---|---|---|
| Swap between LLM servers on request | yes | yes |
| LLMs and ComfyUI share one GPU | – | yes |
| Queue with positions, job states and an event log | – | yes (SQLite + JSONL) |
| Substitution, with the reason reported | – | yes |
| Downloads by Hugging Face repo or GitHub URL | – | yes |
| Interactive GPU sessions for ComfyUI | – | yes |
| Concurrent calls to the resident LLM | proxied | up to `slots`, some reserved for people |
| Where model servers live | processes it launches | systemd units, Docker containers, Proxmox LXCs |

- Only swapping LLMs: use llama-swap.
- One card serving chat *and* diffusion: use this.

## Security

Details: [SECURITY.md](SECURITY.md). Threat model: [docs/threat-model.md](docs/threat-model.md).

- **Token.** A bearer token, compared in constant time.
  - With none set, `serve` won't start.
  - Secrets come from the environment, never the config file.
- **Bind address.** `127.0.0.1` by default. Put TLS in front if you expose it.
- **Allowlist.** Drivers touch only units named in the catalog plus `comfy.unit`.
  - Names, references and paths are validated; files stay under fixed roots; nothing runs through a shell.
- **No SSRF by default.** The broker calls only URLs from its own config and catalog.
  - Fetching a caller's `<slot>_url` is opt-in (`inputs.allow_urls`).
- **Exec recipes are host configuration.** The command comes from the recipe file, never the request.
- **Proxmox: forced command, not a shell.** The host pins the broker's SSH key to one script:

  ```
  command="/usr/local/sbin/gpu-broker-ctl",restrict ssh-ed25519 AAAA... gpu-broker
  ```

  - It accepts only `unit`, `gpu`, `gpustream`, `download`, `comfy-link` and `exec-put` / `exec-info` / `exec-run` / `exec-clean`.
  - Only for the `<container>:<unit>` pairs in `ALLOW_UNITS` (`/etc/gpu-broker-ctl.conf`) and the recipes in `RECIPES`.
  - It re-validates every argument and logs each call. With no config file it allows nothing.
- **Dashboard.** The page carries no data and runs under a strict Content-Security-Policy.
  The token stays in your browser.

## FAQ

| question | answer |
|---|---|
| AMD / ROCm? | Yes. See [docs/hardware.md](docs/hardware.md). |
| Intel, Apple? | Not yet. Starting and stopping servers is vendor-neutral, but there is no GPU reader for them. |
| Multiple GPUs? | Not yet. One broker manages one GPU, the one `gpu.index` selects. |
| Does it run models itself? | No. It controls servers you already run (llama.cpp, vLLM, any OpenAI-compatible server with `/health`, plus ComfyUI). |
| Why not run everything at once? | VRAM. On 24 GB, an 8B LLM at Q4 takes 7–8 GB and a 5B video model at fp16 over 20 GB. Partial offloading makes both slow. |

## Limitations

- **One NVIDIA or AMD GPU.**
  - Per-process VRAM needs the host PID namespace; in a plain container you get totals only.
  - AMD support is tested against fixture files in the kernel's documented formats, not yet a live card.
- **LLM servers** must be OpenAI-compatible and answer `GET /health` with 200.
  - tok/s and time-to-first-token need llama.cpp's `timings` block.
- **Images and video run through ComfyUI only,** with a graph builder per model family in `gpu_broker/templates/`.
  - Some need custom node packs the broker does not install: see [Bundled templates](docs/catalog.md#bundled-templates).
- **One model at a time.** An LLM and a ComfyUI model are never loaded together, even when both would fit.
- **Downloads are not wired into ComfyUI.** Put the files in ComfyUI's model folders yourself.
  - A downloaded model with no template stays `needs_integration`.
- **Docker driver:** existing containers only; needs the root-equivalent Docker socket; no exec recipes.
- **Exec on Proxmox** copies each input file over its own SSH call, as base64.
  - Many `frames` mean many connections; send very large videos as `video_url`.
- **One shared token,** with no per-user accounts or rate limits.

## Roadmap

- Soft preemption and bounded same-model batching for the queue.
- GPU readings for more vendors, and more than one GPU per host.
- Wiring downloaded files into ComfyUI from the API.

## Contributing

Tests never touch a real GPU, host or network:

```bash
git clone https://github.com/emergenthq-net/gpu-broker && cd gpu-broker
python3 -m venv .venv && .venv/bin/pip install -e '.[dev]'
.venv/bin/pytest && .venv/bin/ruff check . && .venv/bin/mypy
```

House rules (layers, no magic values, new graphs and drivers): [CONTRIBUTING.md](CONTRIBUTING.md).

## Changelog

- **0.4.0** (2026-10-03): `gpu-broker setup`, one command from install to a running broker; `gpu-broker init`; docs split into `docs/`.
- **0.3.2** (2026-10-03): the source distribution is self-testing; unpack it, install it with `[dev]`, run `pytest`.
- **0.3.1** (2026-10-03): shipped tests use neutral ids and paths; CI scans every release for private names.

Every release: [CHANGELOG.md](CHANGELOG.md) and [GitHub Releases](https://github.com/emergenthq-net/gpu-broker/releases).

## License

Apache-2.0. See [LICENSE](LICENSE).
