One big model. All your laptops.

Baton splits a model into layers, gives each laptop a share, and passes the work between them over your Wi-Fi.

How it works
This laptop runs (detected, change it for another laptop)
Model
GB
Terminal

        

Sample output of baton serve, two laptops
baton head f1d3bd1e: control on port 7711
reading model metadata: unsloth/Llama-3.2-1B
joined: macbook-air
joined: gaming-pc
benchmarking macbook-air, gaming-pc
Device        Role    Layers
macbook-air   N1      0-5
gaming-pc     N2=Nk   6-15
  macbook-air: 55/55 tensors
  gaming-pc: 92/92 tensors
READY: plan 1 is loaded on 2 node(s)

What happens after you paste

Laptops find each other
No IP addresses to type. Each laptop announces itself on the local network.
Baton measures every laptop
Free memory, GPU, Wi-Fi or cable, and a short speed test.
Each laptop loads its own layers
A faster laptop gets more layers. Each one downloads only its share.

One answer, passed hand to hand

the next word goes back to Laptop A, and the lap starts again

A large model is a tall stack of layers. No single laptop has the memory for all of them, but three laptops together do.

Only a small block of numbers crosses the network on each hop. The model weights never leave the laptop that holds them.

Before you start

What do I need?

Two or more laptops on the same Wi-Fi or the same switch. macOS or Windows. The installer brings its own Python.

Windows asks about the firewall

Choose Allow on private networks. The laptops must be able to reach each other, and Windows blocks that by default.

Does anything leave my network?

The model is downloaded from Hugging Face once. Your prompts and answers stay on your laptops.

Is it fast?

It is slower than one large GPU. The point is to run a model that no single laptop can hold.

Which models work today?

Llama 3.2 and Qwen 2.5 up to 3B at full precision. Larger models need 4-bit weights, which are in progress.