Metadata-Version: 2.4
Name: matmul-mcp
Version: 0.0.1
Summary: MCP server for MatMul: find and launch GPU compute from Claude, Cursor or any MCP client
Author: MatMul
License-Expression: Apache-2.0
Project-URL: Homepage, https://matmul.cloud
Project-URL: Documentation, https://matmul.cloud/docs
Project-URL: Get an API key, https://app.matmul.cloud/app/cli
Project-URL: Python SDK, https://pypi.org/project/matmul-cloud/
Keywords: mcp,model-context-protocol,gpu,cloud,neocloud,claude,cursor
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: mcp<2,>=1.12
Requires-Dist: matmul-cloud>=0.1.0
Dynamic: license-file

# matmul-mcp

An [MCP](https://modelcontextprotocol.io) server for **MatMul**, so you can find and launch GPU
compute from Claude Code, Claude Desktop, Cursor, VS Code or any other MCP client. Ask things like
"find me the cheapest H100 under $3/hr and give me an SSH command", and the assistant drives the
MatMul API for you. Every launch needs your explicit OK on a price quote first.

It runs locally over **stdio** and talks to `https://api.matmul.cloud/v1` with your API key. It is
built on the stdlib-only [`matmul-cloud` Python SDK](https://pypi.org/project/matmul-cloud/) (import name `matmul`) and the official `mcp` Python SDK
(FastMCP).

## Install

You need an API key. Mint one in the console at **https://app.matmul.cloud/app/cli**.

```bash
uvx matmul-mcp                 # run without installing (recommended)
pipx install matmul-mcp        # or install the `matmul-mcp` command
```

`pip install matmul-mcp` also installs the `matmul` CLI (from `matmul-cloud`), so `matmul login`
works too.

Authentication is resolved in this order:

1. the `MATMUL_API_KEY` environment variable;
2. the key saved by `matmul login --api-key <key>` (`~/.config/matmul/config.json`, or
   `$MATMUL_CONFIG`).

| Variable | Meaning |
|---|---|
| `MATMUL_API_KEY` | API key (`Authorization: Bearer ...`). |
| `MATMUL_API_URL` | API base URL. Defaults to `https://api.matmul.cloud`; a trailing `/v1` is accepted. |
| `MATMUL_MCP_MAX_HOURLY_USD` | Optional hard cap: this server refuses to quote or launch any instance above this $/hr, whatever the model asks for. Also caps `find_gpus`. |
| `MATMUL_MCP_ENABLE_JOBS` | `1` adds the preview managed-jobs tools (off by default: see below). |

The product was previously called Lemnos: the old `LEMNOS_*` variable names still work when the
`MATMUL_*` one is unset, and a key saved in `~/.config/lemnos/config.json` is still read.

## Configure your client

In every snippet below, replace `lmk_...` with your key. You can leave `MATMUL_API_KEY` out if
you've run `matmul login` on this machine. The cap is optional but recommended.

### Claude Code

```bash
claude mcp add matmul \
  -e MATMUL_API_KEY=lmk_... \
  -e MATMUL_MCP_MAX_HOURLY_USD=5 \
  -- uvx matmul-mcp
```

Add `--scope user` to make it available in every project.

### Claude Desktop

`~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or
`%APPDATA%\Claude\claude_desktop_config.json` (Windows):

```json
{
  "mcpServers": {
    "matmul": {
      "command": "uvx",
      "args": ["matmul-mcp"],
      "env": {
        "MATMUL_API_KEY": "lmk_...",
        "MATMUL_MCP_MAX_HOURLY_USD": "5"
      }
    }
  }
}
```

If Claude Desktop can't find `uvx`, use its absolute path (`which uvx`).

### Cursor

`~/.cursor/mcp.json` (global) or `.cursor/mcp.json` (per project):

```json
{
  "mcpServers": {
    "matmul": {
      "command": "uvx",
      "args": ["matmul-mcp"],
      "env": {
        "MATMUL_API_KEY": "lmk_...",
        "MATMUL_MCP_MAX_HOURLY_USD": "5"
      }
    }
  }
}
```

### VS Code (GitHub Copilot agent mode)

`.vscode/mcp.json`. VS Code prompts for the key once and stores it securely:

```json
{
  "inputs": [
    { "type": "promptString", "id": "matmul-key", "description": "MatMul API key", "password": true }
  ],
  "servers": {
    "matmul": {
      "type": "stdio",
      "command": "uvx",
      "args": ["matmul-mcp"],
      "env": {
        "MATMUL_API_KEY": "${input:matmul-key}",
        "MATMUL_MCP_MAX_HOURLY_USD": "5"
      }
    }
  }
}
```

## What it exposes

### Tools

| Tool | Does | Annotations |
|---|---|---|
| `find_gpus(gpu?, max_hourly_usd?, region?, count?, limit)` | Live GPU offers, cheapest first | read-only |
| `get_balance(include_ledger?, ledger_limit?)` | Prepaid balance, frozen state, new-account limits, burn rate, runway | read-only |
| `add_credit(amount_usd)` | Returns a **Stripe checkout URL for the human** to open and pay; charges nothing itself | write |
| `list_instances(include_deleted?)` | Instances with status, price, accrued cost, SSH commands | read-only |
| `get_instance(id_or_name)` | One instance; poll it until `running` to get `ssh_command` | read-only |
| `launch_instance(name, gpu, max_hourly_usd, image?, keep_alive?, ssh_key?, count?, region?, confirm_token?)` | Two-step: quote, then launch (see below) | write, spends money |
| `terminate_instance(id_or_name)` | Delete an instance and stop billing | **destructive** |
| `list_ssh_keys()` / `add_ssh_key(name, public_key)` | Manage the keys installed on new machines. `add_ssh_key` refuses anything that looks like a private key. | read-only / write |

With `MATMUL_MCP_ENABLE_JOBS=1`: `run_job`, `list_jobs`, `job_logs`, `cancel_job` (destructive).
These are opt-in because the `/v1/jobs` API is still a preview and has no price guard yet. Turn
them on by default once it has one.

Every read tool is marked `readOnlyHint`, and `terminate_instance` and `cancel_job` are marked
`destructiveHint`, so clients that auto-approve reads still ask before destructive calls.

### Resources and prompts

- `matmul://instances`: your non-deleted instances, as JSON.
- `matmul://balance`: balance and frozen state, as JSON.
- Prompt `gpu_dev_box(gpu, max_hourly_usd, image)`: "launch a GPU dev box". It walks the
  assistant through find, add key, quote, confirm, wait and SSH.

## Money safety

`launch_instance` rents a real machine that bills every hour until it is terminated. The
safeguards:

1. **`max_hourly_usd` is required.** It's sent as the API's `max_hourly_price_cents` price guard.
   The API refuses to launch if the live price has risen above it, and the tool reports "Price
   guard: nothing was launched and nothing was charged".
2. **Two-step confirm (quote → token → launch).** The first call launches nothing. It picks the
   cheapest live offer at or under the ceiling and returns a **quote**: offer, estimated $/hr, the
   ceiling, image or plain VM, SSH key, balance and runway, and a line saying it bills until
   terminated. It also returns a `confirm_token`. Only a second call with that token *and
   identical arguments* launches. The launch result repeats the hourly cost and says to call
   `terminate_instance`.
3. **`MATMUL_MCP_MAX_HOURLY_USD`**, if set, is enforced client-side on both the quote and the
   launch. It lives in the client config, so the model can't change it through the server.
4. **No auto top-up, ever.** `add_credit` only returns a Stripe link for a human to pay. When a
   launch fails because the account is frozen or short on balance, the error points at
   `get_balance` / `add_credit` and states that the server never tops up on its own.

### Why a confirm token, not `dry_run=True`

A `dry_run` flag that defaults to true is only one boolean away from a launch. Nothing stops a
model from sending `dry_run=false` on its first call, and then the user never sees a price. The
token makes the quote a required step:

- **The token proves a quote happened.** It's random (`lq_...`), so the only way to get one is the
  quoting call, and that call's result, with the price, lands in the transcript the user sees.
  The tool description and the server instructions tell the model to get an explicit yes first.
- **It's bound to the request.** Name, GPU, count, region, image, keep-alive, SSH key and price
  ceiling all have to match, and the launch uses the *quoted offer*. So a model can't quote a
  cheap box and launch a different one. A mismatch burns the token.
- **It's single-use and expires after 10 minutes.** A stale "yes" can't be replayed hours later
  at a different price.
- **It works in every MCP client.** It's plain tool calls. It doesn't depend on elicitation
  (which most clients don't support yet) or on the client's per-call approval dialog, which people
  often switch to "always allow". Where clients do show approval prompts, the annotations still
  apply on top.

Tokens are held in the server process's memory. A stdio server is one process per client
session, which is the right lifetime for a quote. A future version could also use MCP
elicitation for the confirm step on clients that support it, and keep the token as the fallback.

## Development

```bash
cd sdk/mcp
uv sync                        # installs mcp + the local ../python SDK (editable path dependency)
uv run pytest                  # unit tests + a real stdio handshake; no network, no real API
uv run matmul-mcp              # run the server on stdio
npx @modelcontextprotocol/inspector uv run matmul-mcp   # poke at it interactively
```

The tests mock the HTTP layer (they patch `urlopen` under the SDK), so each tool runs through
the real SDK code and a real MCP client session. They cover every tool's happy path; error
mapping (401 means run `matmul login` or set `MATMUL_API_KEY`, a 402 or frozen account gives the
balance message, the price guard, 403/404/502/503); the quote → confirm flow, including
single-use, expiry, argument binding and the env cap; and a stdio subprocess
`initialize` + `tools/list` handshake.

To try it against a local API with a fake GPU supplier (no accounts, no spend), use
`apps/api/dev/stage1.sh fake`. It builds and runs the API from a temp dir; never run the API from
`apps/api/`. Then point the server at it:

```bash
PORT=8081 API_KEYS=lmk_dev=dev apps/api/dev/stage1.sh fake       # terminal 1
MATMUL_API_KEY=lmk_dev MATMUL_API_URL=http://localhost:8081 \
  npx @modelcontextprotocol/inspector uv run --project sdk/mcp matmul-mcp   # terminal 2
```

The `matmul-cloud` dependency is a path dependency on `../python` for development (`[tool.uv.sources]`).
Installs from PyPI resolve the plain `matmul-cloud>=0.1.0` requirement from PyPI.

Releases go out through `.github/workflows/release-pypi.yml`: bump `__version__` in
`matmul_mcp/__init__.py`, add a changelog entry below, merge, then push a `matmul-mcp-v<version>`
tag from `main`.

## Changelog

### 0.0.1

First release on PyPI: the stdio server with offers, quote → confirm launch, instances, SSH keys
and billing tools, the optional `MATMUL_MCP_MAX_HOURLY_USD` cap, the preview jobs tools behind
`MATMUL_MCP_ENABLE_JOBS=1`, and `matmul-mcp --version`.

## Remote MCP (hosted, OAuth): later

These are design notes only; none of this is built. The next step is a hosted MCP endpoint, so
people can add MatMul to claude.ai, ChatGPT or Cursor with a URL, with nothing to install and no
key to paste.

- **Endpoint:** `https://mcp.matmul.cloud/mcp`, using the MCP **streamable HTTP** transport. The
  same tool set, served from `FastMCP(...).run("streamable-http")` or ported into the Go API.
  Stateless HTTP mode lets it scale horizontally.
- **Auth:** OAuth 2.1 as the MCP authorization spec describes. The MCP server is a resource server
  and publishes `/.well-known/oauth-protected-resource`, pointing at **WorkOS AuthKit** as the
  authorization server. AuthKit already runs the console login and supports dynamic client
  registration, so MCP clients can register themselves. Access tokens carry the user and org, and
  the server maps them to the same `X-Org-Id`/`X-User-Id` identity the console uses, so RBAC
  (owner/admin/member/viewer) applies per person, the same as in the console.
- **Scopes:** split `instances:read`, `instances:write` (launch/terminate) and `billing:read`,
  with checkout links behind `billing:write`. That way an org can grant read-only access to an
  assistant.
- **Confirm tokens** move from process memory to a shared store (Postgres or KV), keyed per
  user, with the same TTL and binding. Where the client supports **elicitation**, use it for the
  confirm step.
- **Spend limits** move server-side: a per-org "max $/hr via MCP" setting in the console replaces
  the `MATMUL_MCP_MAX_HOURLY_USD` env var.
- **Audit:** log every launch and terminate made over MCP with the OAuth client id, and show the
  log in the console.
