One big model. All your laptops.
Baton splits a model into layers, gives each laptop a share, and passes the work between them over your Wi-Fi.
baton head f1d3bd1e: control on port 7711 reading model metadata: unsloth/Llama-3.2-1B joined: macbook-air joined: gaming-pc benchmarking macbook-air, gaming-pc Device Role Layers macbook-air N1 0-5 gaming-pc N2=Nk 6-15 macbook-air: 55/55 tensors gaming-pc: 92/92 tensors READY: plan 1 is loaded on 2 node(s)
What happens after you paste
- Laptops find each other
- No IP addresses to type. Each laptop announces itself on the local network.
- Baton measures every laptop
- Free memory, GPU, Wi-Fi or cable, and a short speed test.
- Each laptop loads its own layers
- A faster laptop gets more layers. Each one downloads only its share.
One answer, passed hand to hand
Laptop A
layers 0-5turns your words into numbersLaptop B
layers 6-11keeps thinkingLaptop C
layers 12-15picks the next wordA large model is a tall stack of layers. No single laptop has the memory for all of them, but three laptops together do.
Only a small block of numbers crosses the network on each hop. The model weights never leave the laptop that holds them.
Before you start
What do I need?
Two or more laptops on the same Wi-Fi or the same switch. macOS or Windows. The installer brings its own Python.
Windows asks about the firewall
Choose Allow on private networks. The laptops must be able to reach each other, and Windows blocks that by default.
Does anything leave my network?
The model is downloaded from Hugging Face once. Your prompts and answers stay on your laptops.
Is it fast?
It is slower than one large GPU. The point is to run a model that no single laptop can hold.
Which models work today?
Llama 3.2 and Qwen 2.5 up to 3B at full precision. Larger models need 4-bit weights, which are in progress.