Metadata-Version: 2.4
Name: gpumesh
Version: 0.4.0
Summary: Borrow your friends' GPUs: a terminal-based distributed compute mesh in pure Python
License-Expression: MIT
Project-URL: Homepage, https://github.com/Samurai007AK/gpumesh
Project-URL: Documentation, https://github.com/Samurai007AK/gpumesh#readme
Project-URL: Repository, https://github.com/Samurai007AK/gpumesh
Project-URL: Issues, https://github.com/Samurai007AK/gpumesh/issues
Keywords: gpu,distributed,compute,mesh,ml,pytorch,cuda
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: System :: Distributed Computing
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: gpu
Requires-Dist: torch; extra == "gpu"
Provides-Extra: tunnel
Requires-Dist: pyngrok; extra == "tunnel"
Provides-Extra: sysinfo
Requires-Dist: psutil; extra == "sysinfo"
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: build; extra == "dev"
Requires-Dist: twine; extra == "dev"
Dynamic: license-file

# gpumesh

**Borrow your friends' GPUs.** A terminal-based distributed compute mesh written in
pure Python (stdlib only — `torch`, `psutil`, `pyngrok` are optional extras).

You have ML work to run but a weak laptop. Your friend has a strong GPU sitting
idle. `gpumesh` turns any group of machines into a small compute cluster: one
machine coordinates, any number of machines join with a single command, and work
is split between them **proportionally to how fast each machine is**.

```
 your laptop (coordinator)              friend's laptop (worker)
┌───────────────────────┐    HTTP/JSON  ┌────────────────────────┐
│  job queue            │◄─────────────►│  hardware probe        │
│  capability scheduler │   (optional   │  matmul benchmark      │
│  SQLite (WAL)         │    ngrok      │  sandboxed subprocess  │
│  heartbeats + reaper  │    tunnel)    │  executor              │
└──────────▲────────────┘               └────────────────────────┘
           │ submit / status
       client CLI
```

---

## Table of Contents

- [Installation](#installation)
- [Quick Start (5 minutes)](#quick-start-5-minutes)
- [Complete User Guide](#complete-user-guide)
  - [Step 1 — Install](#step-1--install-on-all-machines)
  - [Step 2 — Start Coordinator](#step-2--start-the-coordinator-machine-a)
  - [Step 3 — Join Workers](#step-3--join-workers-machine-b-c-etc)
  - [Step 4 — Submit a Job](#step-4--submit-a-job)
  - [Step 5 — Monitor Progress](#step-5--monitor-progress)
  - [Step 6 — Cancel a Job](#step-6--cancel-a-job)
- [All Commands Reference](#all-commands-reference)
- [Writing Your Own Task Scripts](#writing-your-own-task-scripts)
- [Network Options](#network-options)
- [Security](#security)
- [Troubleshooting](#troubleshooting)
- [How the Scheduler Works](#how-the-scheduler-works)
- [Design Principles](#design-principles)
- [Layout](#layout)
- [Contributing](#contributing)
- [License](#license)

---

## Installation

### Prerequisites

You need **Python 3.9 or newer**. Check your version:

```bash
python --version
```

If you don't have Python 3.9+, download it from [python.org](https://www.python.org/downloads/).

### Install the package

On **every machine** that will participate (coordinator AND workers):

```bash
pip install gpumesh
```

### Optional extras

Depending on your setup, you may want extra packages:

```bash
# For GPU detection and CUDA benchmarks (worker machines with NVIDIA GPUs)
pip install gpumesh[gpu]

# For public internet access via ngrok (coordinator machine)
pip install gpumesh[tunnel]

# For detailed system info like RAM reporting
pip install gpumesh[sysinfo]

# All optional extras at once
pip install gpumesh[gpu,tunnel,sysinfo]
```

### Verify installation

```bash
gpumesh --help
```

You should see:

```
usage: gpumesh [-h] {serve,join,quickjoin,submit,status,cancel,workers} ...

borrow your friends' GPUs

positional arguments:
  {serve,join,quickjoin,submit,status,cancel,workers}
    serve               start a coordinator on this machine
    join                offer this machine's compute to a mesh
    quickjoin           one-click setup: install, detect GPU, and join mesh
    submit              submit a job to the mesh
    status              show job progress and results
    cancel              cancel a running job
    workers             list workers in the mesh
```

---

## Quick Start (5 minutes)

Here is the absolute fastest way to get running with two machines on the same Wi-Fi.

**Machine A — Start the coordinator:**

```bash
gpumesh serve
```

Output:

```
[mesh] coordinator listening on 0.0.0.0:8000
[mesh] token: aB3xK9mN2pQr
[mesh] LAN join command:
       gpumesh join http://192.168.1.10:8000 --token aB3xK9mN2pQr
[mesh] Ctrl+C to stop
```

**Machine B — Join as a worker** (copy the join command from Machine A's output):

```bash
gpumesh join http://192.168.1.10:8000 --token aB3xK9mN2pQr
```

Output:

```
[worker] device=cuda (NVIDIA GeForce RTX 3080) score=85.234 GFLOP/s
[worker] joined mesh as w1
```

**Machine A — Submit a job** (in a new terminal):

```bash
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token aB3xK9mN2pQr --wait
```

That's it! The job runs across all connected workers and results appear when done.

---

## Complete User Guide

This section walks you through everything step by step, with explanations at each stage.

### Step 1 — Install on all machines

Open a terminal on **every machine** you want to use and run:

```bash
pip install gpumesh
```

That's it for basic usage. If a worker machine has an NVIDIA GPU and you want it detected automatically, also run:

```bash
pip install gpumesh[gpu]
```

> **Windows users:** If `pip` doesn't work, try `python -m pip install gpumesh` instead.

> **macOS users:** You may need to use `pip3 install gpumesh` if you have both Python 2 and 3 installed.

### Step 2 — Start the coordinator (Machine A)

The coordinator is the machine that manages the job queue and distributes work. Pick one machine to be the coordinator.

Open a terminal and run:

```bash
gpumesh serve --port 8000 --token YOUR_SECRET_TOKEN
```

Replace `YOUR_SECRET_TOKEN` with any password-like string you want. For example:

```bash
gpumesh serve --port 8000 --token myTeamSecret2026
```

You will see output like:

```
[mesh] coordinator listening on 0.0.0.0:8000
[mesh] token: myTeamSecret2026
[mesh] LAN join command:
       gpumesh join http://192.168.1.10:8000 --token myTeamSecret2026
[mesh] Ctrl+C to stop
```

**Important notes:**

- The `--port` flag chooses which port the server listens on. Default is 8000.
- The `--token` flag sets the password. If you omit it, a random one is generated.
- **Leave this terminal running.** It is the coordinator. Workers connect to it.
- Press `Ctrl+C` to stop the coordinator.

> **Tip:** Write down the LAN join command printed in the output. You will give this to your friends.

#### Optional: Expose publicly via Tailscale (recommended for remote teams)

If your machines are on different networks but all have [Tailscale](https://tailscale.com/download) installed:

```bash
gpumesh serve --port 8000 --token myTeamSecret2026 --tailscale
```

This auto-detects the Tailscale IP and prints it. Workers on the same Tailscale network can connect from anywhere.

#### Optional: Expose publicly via ngrok (for internet access)

If you want anyone on the internet to connect:

```bash
gpumesh serve --port 8000 --token myTeamSecret2026 --public
```

This creates an ngrok tunnel and prints a public URL like `https://abc123.ngrok-free.app`. Requires `pip install gpumesh[tunnel]`.

> **Security warning:** A public tunnel means anyone who guesses your token can run code on your coordinator. Use a strong token and shut down when not in use.

### Step 3 — Join workers (Machine B, C, etc.)

Each machine that contributes compute power runs as a worker.

#### Option A: One-click join (easiest)

```bash
gpumesh quickjoin http://192.168.1.10:8000 --token myTeamSecret2026
```

This command automatically:
1. Checks your Python version
2. Installs gpumesh if missing
3. Detects NVIDIA GPUs
4. Installs PyTorch with CUDA if a GPU is found
5. Benchmarks your hardware
6. Joins the mesh

#### Option B: One-click join with Tailscale

If all machines are on the same Tailscale network:

```bash
gpumesh quickjoin --token myTeamSecret2026 --tailscale
```

This auto-detects the coordinator's Tailscale IP. No URL needed!

#### Option C: Manual join

```bash
gpumesh join http://192.168.1.10:8000 --token myTeamSecret2026
```

#### What happens when you join

You will see output like:

```
[worker] device=cuda (NVIDIA GeForce RTX 3080) score=85.234 GFLOP/s
[worker] joined mesh as w1
```

- **device** — what hardware was detected (`cuda` = NVIDIA GPU, `mps` = Apple Silicon, `cpu` = processor only)
- **device_name** — the specific hardware model
- **score** — a GFLOP/s benchmark number. Higher = faster machine. The scheduler uses this to give heavier tasks to faster machines.
- **worker_id** — your machine's unique ID in the mesh

> **Leave this terminal running!** It is the worker loop. It will:
> - Send heartbeats to the coordinator every 10 seconds
> - Poll for tasks every 2 seconds
> - Execute tasks in sandboxed subprocesses
> - Report results back
>
> Press `Ctrl+C` to leave the mesh gracefully.

#### Optional: Set a per-task timeout

By default, each task has a 240-second (4 minute) wall-clock limit. To change it:

```bash
gpumesh join http://192.168.1.10:8000 --token myTeamSecret2026 --timeout 600
```

This sets the timeout to 600 seconds (10 minutes).

### Step 4 — Submit a job

From any machine (usually the coordinator), open a new terminal and run:

```bash
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token myTeamSecret2026 --wait
```

#### Breaking down the command

| Part | What it does |
|------|-------------|
| `examples/grid_search.py` | The Python script to run for each task |
| `--payloads examples/payloads.json` | A JSON file listing all the task parameters |
| `--url http://192.168.1.10:8000` | The coordinator's address |
| `--token myTeamSecret2026` | The auth token |
| `--wait` | Wait and show progress until the job finishes |

#### Without `--wait` (fire and forget)

```bash
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token myTeamSecret2026
```

This submits the job and immediately returns a job ID:

```
[client] submitted job j_abc123
[client] check progress: gpumesh status j_abc123 --url http://192.168.1.10:8000 --token myTeamSecret2026
```

#### With a job name

```bash
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token myTeamSecret2026 --name "my_experiment" --wait
```

#### Using environment variables (skip --url and --token)

You can set environment variables to avoid typing the URL and token every time:

```bash
export GPUMESH_URL=http://192.168.1.10:8000
export GPUMESH_TOKEN=myTeamSecret2026

# Now you can just run:
gpumesh submit examples/grid_search.py --payloads examples/payloads.json --wait
gpumesh status j_abc123
gpumesh workers
```

### Step 5 — Monitor progress

#### Check which workers are connected

```bash
gpumesh workers --url http://192.168.1.10:8000 --token myTeamSecret2026
```

Output:

```
  w1  GamingPC              cuda  score=85.234    [alive]
  w2  LaptopB               cpu   score=12.156    [alive]
  w3  MacBookAir            mps   score=22.891    [alive]
```

- **alive** means the worker is connected and sending heartbeats.
- **dead** means the worker hasn't sent a heartbeat recently. Its in-progress tasks will be re-queued.

#### Check job status

```bash
gpumesh status j_abc123 --url http://192.168.1.10:8000 --token myTeamSecret2026
```

Output:

```
job: examples/grid_search.py (j_abc123)  finished=False
  task t_001  [done]  cost=1  worker=w1
    result: {"lr": 0.01, "epochs": 100, "l2": 0.0, "val_accuracy": 0.8234}
  task t_002  [done]  cost=2  worker=w3
    result: {"lr": 0.05, "epochs": 200, "l2": 0.0, "val_accuracy": 0.8891}
  task t_003  [running]  cost=2  worker=w2
  task t_004  [pending]  cost=5
  task t_005  [pending]  cost=5
  task t_006  [pending]  cost=10
```

Task statuses: **pending** (waiting), **running** (being executed), **done** (completed successfully), **failed** (error occurred).

### Step 6 — Cancel a job

```bash
gpumesh cancel j_abc123 --url http://192.168.1.10:8000 --token myTeamSecret2026
```

Output:

```
[client] cancelled job j_abc123
  pending tasks cancelled: 3
  running tasks cancelled: 1
  already finished: 2
```

---

## All Commands Reference

### `gpumesh serve`

Start a coordinator on this machine.

```bash
gpumesh serve [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--port PORT` | 8000 | Port to listen on |
| `--token TOKEN` | *(random)* | Auth token. Generated randomly if omitted |
| `--db PATH` | `gpumesh.db` | SQLite database file path |
| `--public` | off | Expose a public URL via ngrok |
| `--tailscale` | off | Auto-detect Tailscale IP for remote access |

**Examples:**

```bash
# Basic: LAN only, random token
gpumesh serve

# Custom port and token
gpumesh serve --port 9000 --token mySecret123

# Public internet via ngrok
gpumesh serve --port 8000 --token mySecret123 --public

# Tailscale network
gpumesh serve --port 8000 --token mySecret123 --tailscale

# Custom database file
gpumesh serve --db /path/to/mydata.db --token mySecret123
```

---

### `gpumesh join`

Offer this machine's compute to a mesh. Runs as a worker until Ctrl+C.

```bash
gpumesh join URL [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--token TOKEN` | `""` | Auth token (or set `GPUMESH_TOKEN` env var) |
| `--timeout SECONDS` | 240 | Per-task wall-clock time limit |

**Examples:**

```bash
# Basic join
gpumesh join http://192.168.1.10:8000 --token mySecret123

# Join with a 10-minute task timeout
gpumesh join http://192.168.1.10:8000 --token mySecret123 --timeout 600

# Using environment variables
export GPUMESH_TOKEN=mySecret123
gpumesh join http://192.168.1.10:8000
```

---

### `gpumesh quickjoin`

One-click setup: install dependencies, detect GPU, benchmark hardware, and join the mesh.

```bash
gpumesh quickjoin [URL] --token TOKEN [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--token TOKEN` | *(required)* | Auth token |
| `--tailscale` | off | Auto-detect coordinator via Tailscale |
| `--port PORT` | 8000 | Coordinator port (used with `--tailscale`) |

**Examples:**

```bash
# Join with explicit URL
gpumesh quickjoin http://192.168.1.10:8000 --token mySecret123

# Join via Tailscale (auto-detects coordinator IP)
gpumesh quickjoin --token mySecret123 --tailscale

# Join via Tailscale on a custom port
gpumesh quickjoin --token mySecret123 --tailscale --port 9000
```

---

### `gpumesh submit`

Submit a job to the mesh. Uploads a Python script and a list of payloads (task parameters).

```bash
gpumesh submit SCRIPT --payloads FILE [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--payloads FILE` | *(required)* | JSON file containing a list of payload objects |
| `--url URL` | `""` | Coordinator URL (or set `GPUMESH_URL` env var) |
| `--token TOKEN` | `""` | Auth token (or set `GPUMESH_TOKEN` env var) |
| `--name NAME` | *(script filename)* | A human-readable name for the job |
| `--wait` | off | Block until the job finishes; show live progress |

**Examples:**

```bash
# Submit and wait for results
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token mySecret123 --wait

# Submit in background
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token mySecret123

# Submit with a custom name
gpumesh submit examples/grid_search.py --payloads examples/payloads.json \
    --url http://192.168.1.10:8000 --token mySecret123 --name "lr_search_v2" --wait
```

---

### `gpumesh status`

Show job progress, task states, and results.

```bash
gpumesh status JOB_ID [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--url URL` | `""` | Coordinator URL (or set `GPUMESH_URL` env var) |
| `--token TOKEN` | `""` | Auth token (or set `GPUMESH_TOKEN` env var) |

**Examples:**

```bash
gpumesh status j_abc123 --url http://192.168.1.10:8000 --token mySecret123
```

---

### `gpumesh cancel`

Cancel a running or pending job. Running tasks are interrupted; pending tasks are cancelled immediately.

```bash
gpumesh cancel JOB_ID [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--url URL` | `""` | Coordinator URL (or set `GPUMESH_URL` env var) |
| `--token TOKEN` | `""` | Auth token (or set `GPUMESH_TOKEN` env var) |

**Examples:**

```bash
gpumesh cancel j_abc123 --url http://192.168.1.10:8000 --token mySecret123
```

---

### `gpumesh workers`

List all workers in the mesh with their hardware info, benchmark scores, and status.

```bash
gpumesh workers [OPTIONS]
```

| Option | Default | Description |
|--------|---------|-------------|
| `--url URL` | `""` | Coordinator URL (or set `GPUMESH_URL` env var) |
| `--token TOKEN` | `""` | Auth token (or set `GPUMESH_TOKEN` env var) |

**Examples:**

```bash
gpumesh workers --url http://192.168.1.10:8000 --token mySecret123
```

---

### `gpumesh --help`

Show help for all commands.

```bash
gpumesh --help
gpumesh serve --help
gpumesh join --help
```

---

## Writing Your Own Task Scripts

The contract between your script and gpumesh is just two lines of glue code:

```python
import json, sys

# 1. Read the payload (parameters) from stdin
payload = json.load(sys.stdin)

# 2. Do your work here
#    - payload is a dict like {"lr": 0.1, "epochs": 200}
#    - os.environ["GPUMESH_DEVICE"] tells you "cuda", "mps", or "cpu"
result = {"answer": 42}

# 3. Print the result as JSON on the LAST line of stdout
print(json.dumps(result))
```

### Complete example

```python
import json
import os
import sys

def main():
    # Read parameters from stdin
    payload = json.load(sys.stdin)

    lr = payload.get("lr", 0.1)
    epochs = payload.get("epochs", 100)
    dataset = payload.get("dataset", "train.csv")

    # Check what device is available
    device = os.environ.get("GPUMESH_DEVICE", "cpu")
    print(f"Running on device: {device}", file=sys.stderr)

    # Do your actual work here...
    accuracy = 0.95  # placeholder

    # Return results as JSON (last stdout line)
    print(json.dumps({
        "accuracy": accuracy,
        "lr": lr,
        "epochs": epochs,
        "dataset": dataset,
    }))

if __name__ == "__main__":
    main()
```

### Payload file format

Create a JSON file with a list of objects. Each object becomes one task:

```json
[
    {"lr": 0.01, "epochs": 100, "cost": 1},
    {"lr": 0.05, "epochs": 200, "cost": 2},
    {"lr": 0.1, "epochs": 500, "cost": 5},
    {"lr": 0.2, "epochs": 1000, "cost": 10}
]
```

### The `cost` field

Each payload can include an optional `"cost"` field (a number, default 1). The scheduler uses cost to distribute work proportionally:

- **Cost 1** = light task → goes to the slowest machine
- **Cost 10** = heavy task → goes to the fastest machine

Without cost, all tasks are treated as equal weight.

### What fits as a task?

Anything **data-parallel** (independent work units):

- Hyperparameter search (grid search, random search)
- Batch inference (process chunks of data)
- Cross-validation folds
- Rendering chunks
- Brute-force search shards
- Data preprocessing batches
- A/B test variants

### Rules

- **Input**: JSON on stdin
- **Output**: JSON on the **last line** of stdout
- **Errors**: Print errors to stderr or exit with a non-zero code
- **Timeout**: Tasks are killed after the timeout (default 240s)
- **Environment**: `GPUMESH_DEVICE` tells you `cuda`, `mps`, or `cpu`

---

## Network Options

| Method | Pros | Cons | Best for |
|--------|------|------|----------|
| **LAN** | No setup, fastest | Same Wi-Fi required | Same room, same network |
| **Tailscale** | Auto-detect, reliable, works remotely | Requires Tailscale account (free) | Teams, remote collaboration |
| **ngrok** | Works from anywhere | Requires ngrok account, slightly slower | Public access, untrusted networks |

### LAN (Same Wi-Fi)

No extra setup needed. Workers connect to the coordinator's local IP address.

```bash
# Machine A (coordinator)
gpumesh serve --port 8000 --token mySecret

# Machine B (worker) — use the LAN IP from Machine A's output
gpumesh join http://192.168.1.10:8000 --token mySecret
```

### Tailscale (Recommended for teams)

1. Install [Tailscale](https://tailscale.com/download) on all machines
2. Log in to the same Tailscale account on all machines
3. Use the `--tailscale` flag:

```bash
# Machine A
gpumesh serve --port 8000 --token mySecret --tailscale

# Machine B (auto-detects coordinator)
gpumesh quickjoin --token mySecret --tailscale
```

### ngrok (Public internet)

1. Install `pip install gpumesh[tunnel]`
2. Create a free [ngrok account](https://ngrok.com) and connect your agent
3. Use the `--public` flag:

```bash
# Machine A
gpumesh serve --port 8000 --token mySecret --public
# prints: [mesh] public URL: https://abc123.ngrok-free.app

# Machine B (anyone with the URL and token)
gpumesh join https://abc123.ngrok-free.app --token mySecret
```

---

## Security

Workers execute code submitted to the coordinator — that is the core functionality. Only share your coordinator URL and token with people you trust.

### Security features

- **Token authentication** — All API requests require a shared secret token
- **Token hashing** — Tokens are hashed with SHA-256 + random salt before storage (never stored in plain text)
- **Rate limiting** — After 5 failed token attempts in 5 minutes, the IP is locked out for 15 minutes
- **IP allowlist** — Optional feature to restrict access to specific IP addresses
- **Subprocess sandboxing** — Each task runs in its own isolated subprocess with:
  - Process group isolation (kills entire tree on timeout)
  - Wall-clock timeout (default 240 seconds)
  - CPU time limits via `RLIMIT_CPU` on POSIX systems

### Security best practices

1. **Use a strong token** — Treat it like a password. Generate one with:
   ```bash
   python -c "import secrets; print(secrets.token_urlsafe(16))"
   ```
2. **Don't leave public tunnels running unattended** — Shut down the coordinator when not in use
3. **Only join meshes you trust** — Workers execute arbitrary code from the coordinator
4. **Use Tailscale for team access** — More secure than public tunnels

---

## Troubleshooting

### Common issues

| Symptom | Cause | Fix |
|---------|-------|-----|
| `ModuleNotFoundError: No module named 'gpumesh'` | Package not installed | Run `pip install gpumesh` |
| `gpumesh: command not found` | Scripts not in PATH | Use `python -m gpumesh.cli` instead |
| `401 bad or missing token` | Token mismatch | Make sure the token matches exactly on coordinator and worker |
| `coordinator unreachable` | Network issue | Check that the coordinator is running and the URL/port is correct |
| `Connection refused` | Coordinator not listening | Make sure `gpumesh serve` is running on the coordinator machine |
| `task timed out after 240s` | Task took too long | Increase timeout with `--timeout 600` on the worker, or split into smaller tasks |
| `task exited with code 1` | Script error | Check your script's stderr output for the error message |
| Worker shows `[dead]` | Worker stopped sending heartbeats | Worker machine may have disconnected or crashed |
| `PyTorch not found` | Optional GPU support | Run `pip install gpumesh[gpu]` for GPU detection |

### Checking connectivity

From the worker machine, test if you can reach the coordinator:

```bash
# Test with curl
curl http://192.168.1.10:8000/api/workers -H "X-Auth-Token: mySecret"

# Or use Python
python -c "import urllib.request; print(urllib.request.urlopen('http://192.168.1.10:8000/api/workers', headers={'X-Auth-Token': 'mySecret'}).read())"
```

### Viewing logs

The coordinator prints status messages to its terminal:
- `[mesh] worker joined:` — A new worker connected
- `[mesh] job submitted:` — A new job was received
- `[mesh] re-queued N task(s)` — Tasks from dead workers were reassigned

Worker terminals show:
- `[worker] running task` — Currently executing a task
- `[worker] task done` — Task completed
- `[worker] task FAILED` — Task errored

### Windows-specific notes

- Process isolation works differently on Windows (no `RLIMIT_CPU`). Tasks are still killed on timeout.
- Use `python -m gpumesh.cli` if `gpumesh` command is not recognized.
- Paths with spaces in the temporary directory may cause issues with some scripts.

---

## How the Scheduler Works

1. **Hardware benchmark** — Every worker benchmarks itself at join time using matrix multiplication. This produces a GFLOP/s **score**.
2. **Cost-based distribution** — Pending tasks are sorted by their `cost` value. When a worker asks for work, the scheduler picks a task whose cost matches the worker's performance percentile among live workers.
3. **Lease system** — Tasks are *leased*, not permanently assigned. If a worker dies (heartbeats stop) or the lease expires, the reaper re-queues the task for someone else.
4. **Bounded retries** — A failed task is retried up to 3 times before being marked as permanently failed.
5. **Pull model** — Workers fetch work from the coordinator (not the other way around). This naturally handles flaky connections and new machines joining.

```
Worker asks for task
        │
        ▼
┌──────────────────┐
│  Scheduler picks │  Fast worker → heavy task
│  matching task   │  Slow worker → light task
└──────────────────┘
        │
        ▼
┌──────────────────┐
│  Task is leased  │  240s timeout, heartbeat required
└──────────────────┘
        │
    ┌───┴───┐
    │       │
  Done    Failed
    │       │
    ▼       ▼
 Result   Re-queue (max 3x)
 posted   then mark failed
```

---

## Design Principles

- **Networking** — Hand-rolled JSON-over-HTTP protocol on `http.server` with no external frameworks. Worker heartbeats, poll-based task leasing that survives flaky connections, shared-token auth, and optional ngrok/Tailscale tunneling for NAT traversal.
- **Database** — SQLite in WAL mode with foreign keys, indexes, and atomic lease acquisition via conditional `UPDATE ... WHERE status='pending'` to prevent duplicate task assignments. Transactions protect multi-row writes.
- **Process isolation** — Every task runs in a fresh subprocess in its own process group. Timeouts kill the entire process tree; POSIX `RLIMIT_CPU` caps runaway loops. The coordinator uses threads with a lock around the shared DB connection.
- **Fault tolerance** — Capability-based scheduling, lease/heartbeat failure detection, idempotent re-queue with bounded retries, and a pull-based model where workers fetch work instead of the coordinator pushing it.
- **Task parallelism** — Independent work units are sharded across machines, not model tensors. This avoids internet latency bottlenecks and matches how production serverless GPU platforms operate.

---

## Layout

```
gpumesh/
  cli.py         command line entry point (serve / join / submit / status / workers)
  server.py      coordinator: threaded HTTP JSON API + lease reaper
  db.py          SQLite layer: workers, jobs, tasks, atomic leasing
  worker.py      agent loop: register, heartbeat, lease, execute, report
  sandbox.py     subprocess isolation: timeouts, process-group kill, rlimits
  capability.py  hardware probe + matmul benchmark -> capability score
  client.py      job submission and status polling
  tunnel.py      optional ngrok public URL
  security.py    token hashing, rate limiting, IP allowlist
examples/
  grid_search.py hyperparameter-search demo task (pure Python)
  payloads.json  six shards with varying cost
```

---

## Contributing

Contributions are welcome! Here is how to get started:

```bash
# Clone the repository
git clone https://github.com/Samurai007AK/gpumesh.git
cd gpumesh

# Install in development mode with dev extras
pip install -e ".[dev]"

# Run the tests
pytest
```

---

## License

MIT License. See [LICENSE](LICENSE) for details.

---

## Links

- **GitHub**: [github.com/Samurai007AK/gpumesh](https://github.com/Samurai007AK/gpumesh)
- **Issues**: [github.com/Samurai007AK/gpumesh/issues](https://github.com/Samurai007AK/gpumesh/issues)
- **PyPI**: [pypi.org/project/gpumesh](https://pypi.org/project/gpumesh)
