Metadata-Version: 2.5
Name: backpropagate
Version: 1.8.3
Summary: Production-ready headless LLM fine-tuning with smart defaults, Windows support, and modular architecture
Project-URL: Homepage, https://github.com/mcp-tool-shop-org/backpropagate
Project-URL: Documentation, https://github.com/mcp-tool-shop-org/backpropagate#readme
Project-URL: Repository, https://github.com/mcp-tool-shop-org/backpropagate.git
Project-URL: Issues, https://github.com/mcp-tool-shop-org/backpropagate/issues
Project-URL: Changelog, https://github.com/mcp-tool-shop-org/backpropagate/blob/main/CHANGELOG.md
Author-email: mcp-tool-shop <64996768+mcp-tool-shop@users.noreply.github.com>
License-Expression: MIT
License-File: LICENSE
Keywords: api,fine-tuning,headless,llm,lora,machine-learning,qlora,training,unsloth
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: accelerate>=0.34.0
Requires-Dist: argcomplete>=3.0.0
Requires-Dist: bitsandbytes>=0.41.0
Requires-Dist: datasets>=2.19.0
Requires-Dist: filelock>=3.0.0
Requires-Dist: packaging>=21.0
Requires-Dist: peft>=0.13.0
Requires-Dist: sentencepiece>=0.2.0
Requires-Dist: tenacity>=8.0.0
Requires-Dist: torch<3,>=2.0.0
Requires-Dist: transformers<6,>=4.46.0
Requires-Dist: trl!=1.1.*,!=1.2.*,!=1.3.*,!=1.4.*,!=1.5.*,<2,>=0.18.0
Provides-Extra: dev
Requires-Dist: bandit>=1.7.0; extra == 'dev'
Requires-Dist: httpx>=0.27.0; extra == 'dev'
Requires-Dist: hypothesis>=6.100.0; extra == 'dev'
Requires-Dist: mutmut>=2.4.0; extra == 'dev'
Requires-Dist: mypy>=1.0.0; extra == 'dev'
Requires-Dist: packaging>=21.0; extra == 'dev'
Requires-Dist: pip-audit>=2.7.0; extra == 'dev'
Requires-Dist: pre-commit>=3.0.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.21.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.0.0; extra == 'dev'
Requires-Dist: pytest-mock>=3.10.0; extra == 'dev'
Requires-Dist: pytest-timeout>=2.3.0; extra == 'dev'
Requires-Dist: pytest-xdist>=3.5.0; extra == 'dev'
Requires-Dist: pytest>=7.0.0; extra == 'dev'
Requires-Dist: ruff>=0.1.0; extra == 'dev'
Provides-Extra: export
Requires-Dist: llama-cpp-python>=0.2.0; extra == 'export'
Provides-Extra: fp8
Requires-Dist: torchao>=0.10.0; extra == 'fp8'
Provides-Extra: full
Requires-Dist: cryptography>=42.0.0; extra == 'full'
Requires-Dist: datasets<4.4.0; extra == 'full'
Requires-Dist: fastapi>=0.115; extra == 'full'
Requires-Dist: llama-cpp-python>=0.2.0; extra == 'full'
Requires-Dist: psutil>=5.9.0; extra == 'full'
Requires-Dist: pydantic-settings>=2.0.0; extra == 'full'
Requires-Dist: pydantic>=2.0.0; extra == 'full'
Requires-Dist: pyjwt>=2.14.0; extra == 'full'
Requires-Dist: reflex>=0.9.2; extra == 'full'
Requires-Dist: structlog>=24.1.0; extra == 'full'
Requires-Dist: torch<2.13; extra == 'full'
Requires-Dist: torchao>=0.10.0; extra == 'full'
Requires-Dist: transformers<=5.5.0; extra == 'full'
Requires-Dist: trl<=0.24.0; extra == 'full'
Requires-Dist: unsloth>=2026.6.8; extra == 'full'
Requires-Dist: wandb>=0.15.0; extra == 'full'
Provides-Extra: full-no-export
Requires-Dist: cryptography>=42.0.0; extra == 'full-no-export'
Requires-Dist: datasets<4.4.0; extra == 'full-no-export'
Requires-Dist: fastapi>=0.115; extra == 'full-no-export'
Requires-Dist: psutil>=5.9.0; extra == 'full-no-export'
Requires-Dist: pydantic-settings>=2.0.0; extra == 'full-no-export'
Requires-Dist: pydantic>=2.0.0; extra == 'full-no-export'
Requires-Dist: pyjwt>=2.14.0; extra == 'full-no-export'
Requires-Dist: reflex>=0.9.2; extra == 'full-no-export'
Requires-Dist: structlog>=24.1.0; extra == 'full-no-export'
Requires-Dist: torch<2.13; extra == 'full-no-export'
Requires-Dist: torchao>=0.10.0; extra == 'full-no-export'
Requires-Dist: transformers<=5.5.0; extra == 'full-no-export'
Requires-Dist: trl<=0.24.0; extra == 'full-no-export'
Requires-Dist: unsloth>=2026.6.8; extra == 'full-no-export'
Requires-Dist: wandb>=0.15.0; extra == 'full-no-export'
Provides-Extra: logging
Requires-Dist: structlog>=24.1.0; extra == 'logging'
Provides-Extra: mlx
Requires-Dist: mlx-lm>=0.31.0; extra == 'mlx'
Provides-Extra: monitoring
Requires-Dist: psutil>=5.9.0; extra == 'monitoring'
Requires-Dist: wandb>=0.15.0; extra == 'monitoring'
Provides-Extra: production
Requires-Dist: cryptography>=42.0.0; extra == 'production'
Requires-Dist: datasets<4.4.0; extra == 'production'
Requires-Dist: fastapi>=0.115; extra == 'production'
Requires-Dist: pydantic-settings>=2.0.0; extra == 'production'
Requires-Dist: pydantic>=2.0.0; extra == 'production'
Requires-Dist: pyjwt>=2.14.0; extra == 'production'
Requires-Dist: reflex>=0.9.2; extra == 'production'
Requires-Dist: structlog>=24.1.0; extra == 'production'
Requires-Dist: torch<2.13; extra == 'production'
Requires-Dist: transformers<=5.5.0; extra == 'production'
Requires-Dist: trl<=0.24.0; extra == 'production'
Requires-Dist: unsloth>=2026.6.8; extra == 'production'
Provides-Extra: security
Requires-Dist: cryptography>=42.0.0; extra == 'security'
Requires-Dist: pyjwt>=2.14.0; extra == 'security'
Provides-Extra: standard
Requires-Dist: datasets<4.4.0; extra == 'standard'
Requires-Dist: fastapi>=0.115; extra == 'standard'
Requires-Dist: reflex>=0.9.2; extra == 'standard'
Requires-Dist: torch<2.13; extra == 'standard'
Requires-Dist: transformers<=5.5.0; extra == 'standard'
Requires-Dist: trl<=0.24.0; extra == 'standard'
Requires-Dist: unsloth>=2026.6.8; extra == 'standard'
Provides-Extra: ui
Requires-Dist: fastapi>=0.115; extra == 'ui'
Requires-Dist: reflex>=0.9.2; extra == 'ui'
Provides-Extra: unsloth
Requires-Dist: datasets<4.4.0; extra == 'unsloth'
Requires-Dist: torch<2.13; extra == 'unsloth'
Requires-Dist: transformers<=5.5.0; extra == 'unsloth'
Requires-Dist: trl<=0.24.0; extra == 'unsloth'
Requires-Dist: unsloth>=2026.6.8; extra == 'unsloth'
Provides-Extra: validation
Requires-Dist: pydantic-settings>=2.0.0; extra == 'validation'
Requires-Dist: pydantic>=2.0.0; extra == 'validation'
Description-Content-Type: text/markdown

<p align="center">
  <a href="README.md">English</a> | <a href="README.ja.md">日本語</a> | <a href="README.zh.md">中文</a> | <a href="README.es.md">Español</a> | <a href="README.fr.md">Français</a> | <a href="README.hi.md">हिन्दी</a> | <a href="README.it.md">Italiano</a> | <a href="README.pt-BR.md">Português (BR)</a>
</p>

<p align="center">
  <img src="https://raw.githubusercontent.com/mcp-tool-shop-org/brand/main/logos/backpropagate/readme.png" alt="Backpropagate" width="400">
</p>

<p align="center">
  <a href="https://github.com/mcp-tool-shop-org/backpropagate/actions/workflows/ci.yml"><img src="https://github.com/mcp-tool-shop-org/backpropagate/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
  <a href="https://pypi.org/project/backpropagate/"><img src="https://img.shields.io/pypi/v/backpropagate" alt="PyPI"></a>
  <a href="https://codecov.io/gh/mcp-tool-shop-org/backpropagate"><img src="https://img.shields.io/codecov/c/github/mcp-tool-shop-org/backpropagate/main" alt="Coverage"></a>
  <a href="https://scorecard.dev/viewer/?uri=github.com/mcp-tool-shop-org/backpropagate"><img src="https://api.scorecard.dev/projects/github.com/mcp-tool-shop-org/backpropagate/badge" alt="OpenSSF Scorecard"></a>
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue" alt="MIT License"></a>
  <a href="https://mcp-tool-shop-org.github.io/backpropagate/"><img src="https://img.shields.io/badge/Landing_Page-live-blue" alt="Landing Page"></a>
</p>

# Fine-tune a 32B QLoRA — or a 7B end to end — on one GPU. Ship it to Ollama.

Backpropagate fine-tunes large language models on a **single** GPU, sized for the card you actually have. Three lines of Python QLoRA a 7B–32B model on one 32 GB consumer card (RTX 5090). One flag, `--full-ft-offload`, full-fine-tunes a 7B-class model by keeping its weights and gradients in host RAM (Linux or WSL2; slow, and measured below). One more command exports to Ollama, then `ollama run` your finetune. Scales down to 16 GB. First-class on Windows. Prefer a browser to Python? `backprop ui` does all of it with no code ([take the tour](https://mcp-tool-shop-org.github.io/backpropagate/handbook/web-ui/)).

```python
from backpropagate import Trainer

trainer = Trainer("Qwen/Qwen2.5-7B-Instruct")
trainer.train("my_data.jsonl", steps=100)
trainer.export("gguf", quantization="q4_k_m")
```

```bash
backprop export ./output --format gguf --quantization q4_k_m --ollama --ollama-name my-model
ollama run my-model
```

That's it. There's no YAML config file. There's no `accelerate launch` ceremony. There's no separate "now convert it to GGUF" tutorial. If you have a CUDA GPU and a JSONL file with your training data, you're three lines away from a working finetune.

## Install

```bash
# Recommended: isolated Python install (no conflicts with system Python or other projects)
pipx install backpropagate

# Or via uv (faster install, same isolation)
uv tool install backpropagate

# Standard pip (if you manage your own virtualenv)
pip install backpropagate
```

If you want the optional features, swap the install for one of these:

```bash
pipx install "backpropagate[standard]"   # adds Unsloth (2x faster training) + the web UI
pipx install "backpropagate[full]"       # adds everything: unsloth, ui, monitoring, export, etc.
```

Prefer Docker? `docker pull ghcr.io/mcp-tool-shop-org/backpropagate:latest` works too. Images ship for both `linux/amd64` and `linux/arm64`, so Apple Silicon and ARM Linux operators get a native image. A canonical `compose.yaml` for "UI in a container" lives at the repo root: put `user:password` in a `ui-auth.txt` next to it, run `docker compose up`, and sign in at `http://127.0.0.1:7860` (the first start builds the frontend, which takes a minute or two). Run history persists in `~/.backpropagate`.

## Where Backpropagate sits in the space

There are several good libraries for fine-tuning LLMs. They're each great at different things:

- **[Axolotl](https://github.com/OpenAccess-AI-Collective/axolotl)** — if you like YAML configs and want a community of recipes to copy from
- **[LLaMA-Factory](https://github.com/hiyouga/LLaMA-Factory)** — if you want DPO/PPO/RLHF and a web GUI
- **[Unsloth](https://github.com/unslothai/unsloth)** — if you need the fastest possible training and you're on a supported model family
- **[torchtune](https://github.com/pytorch/torchtune)** — if you want Meta's first-party PyTorch-native recipes you can edit

Backpropagate is the missing option: **a 3-line Python API for solo operators on a single consumer GPU who want to train an adapter and ship it.** No YAML, no online RL (PPO/GRPO), no multi-node. There is a browser UI for the same loop if you would rather not write code. Just the loop everyone actually needs and the export step that gets in the way.

If you tried one of the libraries above and bounced off the config-file ceremony, or hit a model-family gap, or wanted Windows-first defaults — Backpropagate is for you.

## What you can fine-tune on one GPU

Backpropagate sizes the run to your card. These are **measured** numbers on a 32 GB RTX 5090: the QLoRA rows from 2026-10-03 (receipts: [`docs/receipts/2026-10-03-presets/`](docs/receipts/2026-10-03-presets/)), the full fine-tuning rows from 2026-09-30 (receipts: [`docs/receipts/2026-09-30-offload/`](docs/receipts/2026-09-30-offload/)). QLoRA peaks are at the preset's full context window with batch 1, which is the worst case for that preset; shorter examples use less.

| Model | Method | Measured on a 32 GB card |
|---|---|---|
| **14B** (Qwen2.5-14B) | QLoRA | **18.7 GiB** peak at 4096 context (20.0 GiB reserved). |
| 24B (Mistral-Small-24B) | QLoRA | 22.8 GiB peak at 4096 context (24.2 GiB reserved). |
| **32B** (Qwen2.5-32B) | QLoRA | **Fits:** 26.0 GiB peak at 2048 context (27.2 GiB reserved, about 4 GiB to spare). |
| 3B | `mode="full"` (true full fine-tuning, on the GPU) | **22.0 GiB** peak (system-wide), 0.30 s/step at batch 4, 512 context. 7.5 GiB of that is paged optimizer state, which can spill to host RAM on a smaller card (untested). |
| **7B-class** (Qwen2.5-7B, 7.6B params) | `mode="full" --full-ft-offload` | **Trains:** 5.3 GiB VRAM, **30.8 GiB host RAM** (32.2 GiB while saving), **14.7 s/step**. Linux or WSL2 only. |

Not re-measured in that session: 7B QLoRA, Llama-3.1-8B (gated repository, no token on the test machine), and pure-GPU full fine-tuning above 3B. Figures for those elsewhere in the docs are estimates.

Two things most single-GPU libraries send you elsewhere for, **24–32B QLoRA** and **single-card 7B-class full fine-tuning**, Backpropagate does on one consumer card, then exports the result straight to Ollama.

**Full fine-tuning has two paths.** Without offload, the model, its gradients and the optimizer state all sit on the GPU. The library caps model size by detected VRAM (**16 GB → 4B, 24 GB → 5B, 32 GB → 6B**); those caps come from memory arithmetic and are measured only up to 3B. Override with `--full-ft-ceiling-billions`.

`--full-ft-offload` keeps weights and gradients in host RAM and streams them to the GPU (FSDP2 CPU offload). What it costs, measured:

- **Host RAM:** the fit check asks for about 3.7 GiB per billion parameters plus 10 GiB, which is conservative (39 GiB at 7.6B against the 32.2 GiB measured). The run is refused up front, with the numbers, if the machine cannot hold it. A 7.6B model does not fit under a 28 GB WSL2 memory cap; about 5B is the practical limit there.
- **Speed:** 14.7 s/step at 7.6B (batch 1) and 5.1 s/step at 3B (batch 4), against about 0.63 s/step for 3B on the GPU at batch 4. Use it only when the model does not fit without it. A faster version is planned.
- **Optimizer:** Adafactor, not AdamW. Weights stay in bf16 and each update is written back with stochastic rounding; there is no fp32 copy.
- **Quality:** on one 3B run (150 steps, held-out loss, one seed) it reached about 85% of the improvement that ordinary full fine-tuning got (2.45 → 1.93 against 2.45 → 1.84). One seed is not a benchmark.
- **Scope:** plain supervised fine-tuning. No packing, no response-only masking, no intermediate checkpoints, no resume. Linux or WSL2 only (FSDP2 needs NCCL); on Windows-native it stops with `DEP_FSDP_UNAVAILABLE`.
- **Not yet tested:** long runs, gradient accumulation above 1, and a physical 64 GB machine (the test machine had more RAM, with a 60 GiB limit enforced by the test).

A model that does not fit exits with `RUNTIME_FULL_FT_MODEL_TOO_LARGE` and names the way out. See [the full fine-tuning handbook page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/full-fine-tuning/).

### Scales down to 16 GB

The 16 GB envelope (RTX 4080 / 5080 / 4070 Ti Super) is still first-class: 7B QLoRA (the adapter size is picked to fit: rank 64 on a 16 GB card, where rank 256 needs about 17 GB), and true full fine-tuning of a ~3B model (SmolLM3-3B, Qwen2.5-3B, Llama-3.2-3B/1B) via `mode="full"`, (22.0 GiB measured on a 32 GB card at 3B, of which 7.5 GiB is paged optimizer state that can spill to host RAM; whether that runs acceptably on a 16 GB card has not been tested). With `--full-ft-offload` the GPU holds far less: with VRAM capped on the test card, a 3B model trained under a 6 GiB cap and 4B and 7.6B models under an 8 GiB cap. Those are emulated caps on a 32 GB card, not runs on real 8 GB hardware. The same code picks the batch size and ceiling that fit whatever card it detects.

2-bit quantization (AQLM / QuIP#) stays **out of scope** — a 2-bit base can't be cleanly merged back into full-precision weights, which breaks the mergeable-adapter → GGUF → Ollama export contract (the whole point of the pipeline). The headroom levers Backpropagate ships instead — QLoRA, `mode="full"`, `--full-ft-offload`, and the FP8 compute path (`--fp8`, Blackwell/Hopper) — all stay mergeable and exportable.

## What Backpropagate is NOT for

If your use case is below, you'll have a better time with a different library — Backpropagate is not the right pick and trying to make it work would cost more than just reaching for the right tool. Reading this section before you start saves the install-and-bounce cycle:

- **Full-parameter fine-tuning of 13B+ models** — Backpropagate full-fine-tunes up to about 6B on a 32 GB GPU and a 7B-class model with `--full-ft-offload` (see [the envelope](#what-you-can-fine-tune-on-one-gpu)). A full fine-tune of a 13B+ model wants multi-GPU FSDP or a bigger card. Before spending that compute, weigh the evidence both ways. [Thinking Machines 2025](https://thinkingmachines.ai/blog/lora/) reports LoRA matching full fine-tuning when it is applied to every layer and the dataset fits the adapter's capacity, at about two-thirds of the compute per pass. [Biderman et al. 2024](https://arxiv.org/abs/2405.09673) found that in standard low-rank settings LoRA substantially underperforms full fine-tuning on code and math, while forgetting less. For instruction-following, persona and style work on modest datasets, QLoRA up to 32B is usually the better use of one card.
- **Online RL — PPO / GRPO / RLVR** — Backpropagate does single-stage SFT plus reference-free preference tuning (ORPO in v1.5; SimPO + KTO in v1.6). What it does *not* do is online reinforcement learning — PPO, GRPO, or RLVR — which needs a reward model or a generation-and-scoring loop on top of the training step. For those, use TRL directly or LLaMA-Factory. (Reference-free preference tuning fits the single-stage envelope because there's no separate reference model to hold in memory; see the ORPO note under [Quick Start](#quick-start).)
- **Multi-node training** — single GPU on one machine only. Multi-GPU on one machine works (via `accelerate launch`) but isn't officially supported.
- **macOS training on the CUDA rail** — Apple Silicon doesn't have CUDA, so the CUDA path runs on a Linux or Windows box with an NVIDIA GPU. You can still run the trained model on a Mac via Ollama. An **experimental, unverified-preview** MLX rail (`--backend mlx`) trains a LoRA adapter natively on Apple Silicon — see [Apple Silicon (MLX)](#apple-silicon-mlx--unverified-preview). It is LoRA-SFT-only and **not dogfood-verified on real silicon** (no support), so for anything beyond a LoRA SFT (ORPO, full fine-tune, FP8, multi-run) you want the CUDA rail.
- **Anything outside the tested model families** — Qwen 2.5 / 3.5 (7B / 4B), Phi-4-mini-3.8B, SmolLM3-3B, Llama 3.2 (3B / 1B), Mistral 7B. Other models often work but aren't pinned in CI.

If you need any of those things, reach for one of the libraries listed above. They're better at them.

## What Backpropagate gives you

Four things, in one install:

**1. A real 3-line API that runs without a config file.**
The snippet at the top of this README runs end-to-end. No `accelerate config`, no YAML, no Hydra overrides. Just `Trainer(model).train(data)` and you have a finetune.

**2. Windows that actually works.**
Most ML libraries treat Windows like an afterthought. Backpropagate is developed and tested on Windows 11 with RTX 50-series cards. The library handles the runtime quirks for you — it knows how to pre-tokenize your data so Windows multiprocessing doesn't crash, it automatically disables xformers on RTX 40/50 cards where it would break, and it picks dataloader settings that don't blow up. You don't have to know any of this. It just runs.

**3. Built for unattended runs.**
Training takes hours. You don't want to babysit it. Backpropagate is designed to be left running:

- If you run out of GPU memory, it automatically halves the batch size and retries — up to three times. No hand-tuning.
- If your GPU gets too hot, it pauses until things cool down and then continues.
- Every checkpoint is written atomically — if your laptop crashes mid-save, the previous good checkpoint is still intact.
- Every training run gets a unique ID that's stamped onto every log line, every checkpoint, and every Weights & Biases entry. If something goes wrong, one ID lets a maintainer correlate everything.
- Errors come with stable codes (`RUNTIME_GPU_OOM`, `DEP_OLLAMA_REGISTRATION_FAILED`, etc.) so you can grep your logs and the [troubleshooting guide](https://mcp-tool-shop-org.github.io/backpropagate/handbook/troubleshooting/) for the fix. CUDA-specific failures have a dedicated [CUDA troubleshooting page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/troubleshooting-cuda/).

**4. One command from trained adapter to `ollama run`.**
Lots of libraries train a model. Few of them get out of your way when you want to actually use it. Backpropagate exports to GGUF (the format Ollama uses) and registers an Ollama model in one command. You go from "training done" to "I can chat with my finetune" in about 30 seconds.

## Quick Start

From the command line, with a 5-conversation example dataset:

```bash
pipx install "backpropagate[standard]"
curl -LO https://raw.githubusercontent.com/mcp-tool-shop-org/backpropagate/main/examples/quickstart.jsonl

backprop train --data quickstart.jsonl --model Qwen/Qwen2.5-7B-Instruct --steps 10
backprop generate ./output "What is Python?"      # did it learn anything?
backprop export ./output --format gguf --quantization q4_k_m --ollama --ollama-name my-first-finetune
ollama run my-first-finetune
```

`backprop train` writes the adapter to `./output` (change it with `--output`). In Python the same thing is:

```python
from backpropagate import Trainer

trainer = Trainer("Qwen/Qwen2.5-7B-Instruct")
trainer.train("quickstart.jsonl", steps=10)
trainer.export("gguf", quantization="q4_k_m")
```

Use a virtual environment with `pip install "backpropagate[standard]"` for the Python API; `pipx` installs the `backprop` command in its own environment, so `import backpropagate` will not find it.

**What GGUF export needs.** The export merges your adapter into the base model and converts it with llama.cpp's converter script. You need either:

- a llama.cpp **source checkout** (`git clone https://github.com/ggml-org/llama.cpp ~/llama.cpp`) plus `pip install sentencepiece protobuf` in the same environment, or
- Unsloth with its own llama.cpp already built.

With `--ollama`, the `q4_k_m` quantization is done by `ollama create`, so nothing has to be compiled. Backpropagate never lets Unsloth install system packages to build llama.cpp for you; set `BACKPROPAGATE_UNSLOTH_AUTO_INSTALL=1` if you want that. Details: [export](https://mcp-tool-shop-org.github.io/backpropagate/handbook/export/).

For your own data, format your JSONL one example per line:

```jsonl
{"conversations": [{"from": "human", "value": "What is Python?"}, {"from": "gpt", "value": "A programming language."}]}
{"conversations": [{"from": "human", "value": "Explain recursion."}, {"from": "gpt", "value": "A function that calls itself."}]}
```

Alpaca (`instruction` / `output`), OpenAI chat (`messages`), and raw text formats also work — Backpropagate auto-detects the format.

### The loop: check the data, train, evaluate, export

```bash
backprop data report my_data.jsonl                     # duplicates, length outliers, format problems
backprop data split my_data.jsonl --heldout-ratio 0.1  # a held-out set the model never trains on
backprop train --data my_data.train.jsonl --steps 200 --output ./run-a
backprop eval <run-id> --heldout my_data.heldout.jsonl # held-out loss + sample generations
backprop eval <run-b> --vs <run-a>                     # did the change help?
backprop export ./run-a --format gguf --ollama --ollama-name my-model
```

Evaluation is judge-free by design: held-out loss plus deterministic task metrics (`normalized_exact_match`, `token_f1`, `contains`, `regex`, `pass_rate`). To use an LLM judge, run it over `backprop generate` output yourself. See [recipes](https://mcp-tool-shop-org.github.io/backpropagate/handbook/recipes/).

### Preference tuning (ORPO, SimPO, KTO)

Train on preferences instead of plain demonstrations. ORPO is reference-free and single-stage — it folds the preference signal into the SFT step, so there's no separate reward or reference model and the 3-line shape is unchanged. Pass `--method orpo` (CLI) or `method="orpo"` (Python) and feed it a dataset of `{prompt, chosen, rejected}` (or just `{chosen, rejected}`) rows:

```jsonl
{"prompt": "What is Python?", "chosen": "A high-level programming language known for readability.", "rejected": "idk look it up"}
{"prompt": "Explain recursion.", "chosen": "A function that calls itself with a smaller input until a base case.", "rejected": "when something repeats"}
```

```python
from backpropagate import Trainer

trainer = Trainer("Qwen/Qwen2.5-7B-Instruct", method="orpo")
trainer.train("preferences.jsonl", steps=100)
trainer.export("gguf", quantization="q4_k_m")
```

```bash
backprop train --data preferences.jsonl --method orpo --steps 100
```

The default learning rate auto-lowers to `8e-6` for ORPO (the loss is sharper than plain SFT); tune `--orpo-beta` (default `0.1`) to weight the odds-ratio penalty. ORPO is `mode="lora"` only.

**New in v1.6 — SimPO and KTO.** `--method simpo` ([Meng et al. 2024](https://arxiv.org/abs/2405.14734)) is reference-free with a length-normalized reward and takes the same paired `{prompt, chosen, rejected}` data as ORPO (`--simpo-beta`, `--simpo-gamma`). `--method kto` ([Ethayarajh et al. 2024](https://arxiv.org/abs/2402.01306)) takes **unpaired** `{prompt, completion, label}` data — per-example thumbs-up/down — for the large class of feedback that isn't curated A/B pairs; it auto-balances the desirable/undesirable loss weights from your label counts. Both are `mode="lora"` only and stay in the single-GPU SFT envelope (no separate reference model). See the [preference-tuning handbook](https://mcp-tool-shop-org.github.io/backpropagate/handbook/preference-tuning/) for which to use. For online RL (PPO/GRPO) see [What Backpropagate is NOT for](#what-backpropagate-is-not-for).

### Reasoning-trace SFT (R1 distillation)

Distill a reasoning model the easy way. Pass `--reasoning-trace` (CLI) or `Trainer(..., reasoning_trace=True)` (Python) and feed it traces that keep a `<think>...</think>` chain-of-thought inside the assistant turn — the pure-SFT half of [DeepSeek-R1](https://arxiv.org/abs/2501.12948) distillation, no RL required. Backpropagate keeps `<think>` in the training target, drops empty / over-long traces (trace-length filtering), and raises the default `max_seq_length` to 8192 for the longer CoT. Critically, `<think>` stays **plain text** — no special tokens, no embedding resize — so the merged GGUF still exports to Ollama like any other fine-tune. SFT only. See the [reasoning-trace recipe](https://mcp-tool-shop-org.github.io/backpropagate/handbook/recipes/#reasoning-trace-sft-r1-distillation) for the dataset shape and the tunable token band.

### Apple Silicon (MLX) — unverified preview

> ⚠️ **Unverified preview — not part of the supported feature set.** The MLX rail is built and unit-tested but has **not** been dogfood-verified on real Apple Silicon (`mlx-lm` is Apple-only and can't run on the NVIDIA rigs Backpropagate is developed on). Treat everything below as experimental, use at your own risk, and [report anomalies](#reporting-bugs) if you run it on an M-series Mac.

**One API, two rails.** CUDA is the canonical, verified backend; MLX is a second rail that trains on an M-series Mac via Apple's [`mlx_lm.lora`](https://github.com/ml-explore/mlx-lm) toolchain (unified memory, no CUDA). The 3-line shape picks the rail by hardware — `backend='auto'` (the default) routes to CUDA on NVIDIA and to MLX on Apple Silicon, so existing CUDA rigs are byte-identical:

```python
from backpropagate import Trainer

# On an M-series Mac with `pip install 'backpropagate[mlx]'`:
trainer = Trainer("mlx-community/Qwen2.5-0.5B-Instruct-4bit", backend="mlx")
trainer.train("examples/quickstart.jsonl", steps=100)
```

```bash
backprop train --data my_data.jsonl --backend mlx --steps 100
```

The MLX rail is **LoRA SFT only** — no ORPO, no FP8, no `mode='full'`, no multi-run (each is rejected with `CONFIG_INVALID_SETTING`; use `backend='cuda'`/`'auto'` on an NVIDIA box for those). The resulting adapter is plain safetensors and exports to Ollama through the same path as the CUDA rail.

> Forcing `--backend mlx` on a non-Apple host errors with `CONFIG_INVALID_SETTING`; a missing `mlx_lm` toolchain on a Mac raises `DEP_MLX_UNAVAILABLE`.

For more end-to-end workflows (fine-tune-and-push-to-HF-Hub, resume after OOM, multi-run SLAO across a long campaign, etc.) see the [handbook recipes page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/recipes/).

### Web UI (optional)

If you'd rather click than type Python, install the UI extra and launch:

```bash
pipx install "backpropagate[ui]"
backprop ui --port 7862
```

Open the URL it prints, `http://127.0.0.1:7862/?token=...` (each launch makes a new token; the first start builds the frontend and can take a minute or two). It is a local web interface for training: start a run, a multi-run sweep or an export, watch it live (steps, loss, time left, GPU temperature and memory), and stop it with a saved checkpoint. Each job runs in its own process, one at a time, and a page reload picks a running job back up. The Dataset page shows what a file contains, saves a cleaned copy (repeats and empty examples removed) and hands it to the training form. Past runs and the models in your Hugging Face cache have their own pages, and every setting has an "i" that explains it. The [web UI tour](https://mcp-tool-shop-org.github.io/backpropagate/handbook/web-ui/) shows each page. The UI is local-only by default. To expose it to other devices, see [Web UI](#web-ui) below for the `--share` + `--auth` security contract.

## Multi-run training

If you want to fine-tune incrementally across multiple datasets — say you get new training data every week and want to add it without forgetting what you learned before — Backpropagate's `multi_run` mode is for you:

```python
from backpropagate import Trainer

trainer = Trainer("Qwen/Qwen2.5-7B-Instruct")

result = trainer.multi_run(
    dataset="HuggingFaceH4/ultrachat_200k",
    num_runs=5,
    steps_per_run=100,
    samples_per_run=1000,
)
```

This runs five training passes, merging the adapter between runs in a way that preserves earlier knowledge while incorporating new examples. The technique is based on recent continual-learning research — see [References](#references) at the bottom of this README.

The CLI version:

```bash
backprop multi-run --data my_data.jsonl --runs 5 --steps 100 --samples 1000
```

## Resume from checkpoint

A 5-run training that crashes at run 4 is recoverable. Every multi-run session writes its run ID into the on-disk history and checkpoint manifest, so picking up where you left off is one command:

```bash
backprop resume <run-id>
backprop multi-run --data ... --resume <run-id>
backprop train --data ... --resume <run-id>     # single-run resume
```

The default behavior of `backprop multi-run` (no `--resume`) auto-detects an in-progress entry in the same output directory and continues it. To force a clean start, point at a fresh output directory.

## Training history

Every `backprop train` and `backprop multi-run` invocation records a row in `<output>/run_history.json` — model used, dataset, hyperparameters, status, final loss, loss history. You can list and inspect past runs:

```bash
backprop list-runs                         # last 20 runs
backprop list-runs --status failed         # filter by status
backprop list-runs --json --limit 100      # machine-readable
backprop show-run abcd1234                 # detail view (partial ID is fine)
```

## Experiment tracking

Backpropagate auto-detects installed experiment trackers (Weights & Biases, TensorBoard, MLflow) and wires them in. If `wandb` is installed and you're logged in, every run automatically logs to W&B with a run name that matches the on-disk run ID — so you can grep across W&B, your logs, and `run_history.json` using one identifier.

```bash
pip install backpropagate[monitoring]   # installs wandb + psutil
wandb login                             # one-time setup
backprop train --data my_data.jsonl
```

Override with `Trainer(report_to=["wandb"])`, `Trainer(report_to=["tensorboard"])`, or `Trainer(report_to="none")` to opt out.

## Web UI

The Reflex web interface is opt-in — install with `pipx install "backpropagate[ui]"` and launch:

```bash
backprop ui --port 7862
```

The UI runs locally: open the URL it prints, `http://127.0.0.1:7862/?token=...`. Without `--auth`, every launch generates a new token and the UI refuses requests that lack it. From it you can look at and clean a dataset, train (a single run or a multi-run), watch the run live, stop it with a saved checkpoint, and export the result. Each job runs in its own process, one at a time, and closing the UI stops it. The [web UI tour](https://mcp-tool-shop-org.github.io/backpropagate/handbook/web-ui/) walks through every page with screenshots.

To expose it to other devices (other people on your network, a public URL, etc.) you must pair `--share` (or `--host`) with `--auth`:

```bash
backprop ui --share --auth alice:hunter2
```

`backprop ui --share` without `--auth` exits with an error. The reason: `--share` publishes a URL anyone on the internet can reach, and without authentication that means anyone can drive your training pipeline and read your HuggingFace token. There is no opt-out for this — if you don't want to set credentials, use SSH port-forwarding instead:

```bash
# On the client:
ssh -L 7862:localhost:7862 <your-training-host>
# On the server:
backprop ui                             # no --share
# Then open the URL the server printed (http://127.0.0.1:7862/?token=...) locally
```

See [handbook/security.md](https://mcp-tool-shop-org.github.io/backpropagate/handbook/security/) for the full threat model.

Filesystem writes from the UI are sandboxed to a single directory:

- Default: `~/.backpropagate/ui-outputs`
- Override: set `BACKPROPAGATE_UI__OUTPUT_DIR=/path/you/own`
- The override is denylist-validated — system or credential paths (`/etc`, `~/.ssh`, `~/.aws`, `C:\Windows\System32`, etc.) are refused

## Platform notes

**Requirements:** Python 3.10+ · NVIDIA GPU with CUDA · PyTorch 2.0+. An 8 GB card trains the 1B to 3B presets, 16 GB a 7B, and 32 GB up to 32B with QLoRA.

Python 3.10 is supported through at least v1.6; it reaches upstream end-of-life in October 2026 and is scheduled for removal in the first release after that. For new installs, prefer Python 3.11 or 3.12 — 3.11 is the most-tested floor.

Backpropagate handles the runtime quirks of training on different platforms, but it can't fix install-time problems. The two most common are:

- **Wrong CUDA wheel.** PyTorch is published one binary per CUDA version. If you pick the wrong one, you silently get CPU-only PyTorch and training is impossibly slow. Use the wheel picker at <https://pytorch.org/get-started/locally/> for your driver. Run `nvidia-smi` to see your driver / CUDA version.
- **Windows + GGUF export.** The `[export]` extra builds `llama-cpp-python` from source, which needs Visual Studio Build Tools (C++ component) and CMake.

**macOS:** the CUDA rail is not supported (no CUDA) — a CUDA-routed `trainer.train()` raises `DEP_GPU_NOT_AVAILABLE`, and you can run the trained adapter on a Mac via Ollama. An **experimental, unverified-preview** MLX rail (`--backend mlx`, `pip install 'backpropagate[mlx]'`) trains a LoRA adapter natively on Apple Silicon via `mlx_lm.lora` — LoRA SFT only, and **not dogfood-verified on real silicon** (see [Apple Silicon (MLX)](#apple-silicon-mlx--unverified-preview)). For the CUDA path, or for ORPO / full fine-tune / FP8 / multi-run, use a CUDA Linux or Windows machine.

See the [troubleshooting handbook page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/troubleshooting/) for the long-form install fix-it guide, and the dedicated [CUDA troubleshooting page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/troubleshooting-cuda/) for driver / VRAM / xformers / bf16-vs-fp16 issues.

## CLI

Every Python API has a CLI mirror:

```bash
backprop train --data my_data.jsonl --model Qwen/Qwen2.5-7B-Instruct --steps 100
backprop multi-run --data my_data.jsonl --runs 5 --steps 100
backprop export ./output --format gguf --quantization q4_k_m --ollama --ollama-name my-model
backprop ui --port 7862
backprop info                          # environment + version snapshot
backprop list-runs                     # past training runs
backprop show-run <run-id>             # detail view
backprop resume <run-id>               # resume a crashed run
backprop push ./output/lora --repo me/my-model    # push adapter to HuggingFace Hub
backprop diff-runs <run-a> <run-b>     # diff two runs side by side
backprop replay <run-id>               # re-run with same config / dataset
backprop export-runs --format jsonl    # bulk export run history
```

Full reference at [the CLI handbook page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/cli-reference/), or `backprop <subcommand> --help`.

## Configuration

Every setting can be overridden with an environment variable using the `BACKPROPAGATE_` prefix:

| Variable | Default | Notes |
|---|---|---|
| `BACKPROPAGATE_LOG_LEVEL` | `INFO` | `DEBUG` / `INFO` / `WARNING` / `ERROR` |
| `BACKPROPAGATE_LOG_JSON` | auto | Force JSON or console logs |
| `BACKPROPAGATE_MODEL__NAME` | `Qwen/Qwen2.5-7B-Instruct` | Default model |
| `BACKPROPAGATE_TRAINING__LEARNING_RATE` | `2e-4` | Learning rate |
| `BACKPROPAGATE_LORA__R` | `256` | LoRA rank. Setting it turns the automatic adapter-size choice off (see `--lora-preset` under [Model presets](#model-presets)). |
| `BACKPROPAGATE_UI__OUTPUT_DIR` | `~/.backpropagate/ui-outputs` | UI filesystem sandbox |

Nested keys use double underscore (`MODEL__NAME`, not `MODEL_NAME`). The full reference is at [the env-vars handbook page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/env-vars/).

## Model presets

| Preset | GPU memory | License | Notes |
|---|---|---|---|
| Qwen-3.5-4B | 6 / 7 / 11 GB | Apache 2.0 | Recommended default for sub-5B. Best quality at this size. |
| Phi-4-mini-3.8B | 6 / 7 / 12 GB | MIT | Strong on reasoning / math / code. Strict license-clean. |
| SmolLM3-3B | 4 / 5 / 10 GB | Apache 2.0 | Fully open recipe. Native 64K context. |
| Qwen 2.5 7B | 9 / 11 / 17 GB | Apache 2.0 | Existing default. Best quality of the legacy 7B presets. |
| Qwen 2.5 3B | 4 / 6 / 10 GB | Qwen-Research | ⚠ research license — see Qwen license terms before commercial use. |
| Llama 3.2 3B | 4 / 6 / 9 GB | Llama Community | Solid alternative to Qwen 3B with permissive caveats. |
| Llama 3.2 1B | 2 / 3 / 5 GB | Llama Community | For quick experiments on small cards. |
| Mistral 7B | 6 / 8 / 14 GB | Apache 2.0 | Comparable to Qwen 7B, different chat template. |
| Llama-3.1-8B | 9 / 11 / 18 GB | Llama-3.1-Community | 8B QLoRA, 128K native context (the >700M-MAU clause needs a separate Meta license). |
| **Qwen2.5-14B** | 18.7 GiB peak at 4096 ctx (QLoRA) | Apache 2.0 | **The 32 GB daily driver.** rank/alpha 32, 8-bit AdamW. The 4-bit weights alone are about 8.5 GB; a full 4096-token window needs the rest. |
| Mistral-Small-24B | 22.8 GiB peak at 4096 ctx (QLoRA) | Apache 2.0 | 24B QLoRA on a 32 GB card. The 4-bit weights alone are about 18 GB. |
| **Qwen2.5-32B** | 26.0 GiB peak at 2048 ctx (QLoRA) | Apache 2.0 | **Top of the 32 GB envelope.** Fits at `max_len 2048` with 8-bit AdamW. |

Other models often work; the rows above are the curated presets — the 14B–32B tier is QLoRA-tuned for a 32 GB card (the measured envelope). For the presets up to 8B, the three figures are QLoRA estimates for the `fast`, `balanced` and `quality` adapter sizes at batch 1 with 2,048-token examples; they err on the high side, and shorter examples use less. The adapter size is chosen for your card: `--lora-preset auto` (the default) takes the largest of `quality` (rank 256 on every linear layer, per Biderman 2024 and Thinking Machines 2025), `balanced` (rank 64 on every linear layer) and `fast` (rank 16 on two layers per block) that fits the memory free on your GPU. Name one to force it. `backprop estimate-vram` prints the estimate for any model and settings.

## Troubleshooting

A short index of the most common first-run failures. The full reverse index is at [the troubleshooting handbook page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/troubleshooting/). For driver / VRAM / mixed-precision deep-dive see the [CUDA troubleshooting page](https://mcp-tool-shop-org.github.io/backpropagate/handbook/troubleshooting-cuda/).

| Symptom | Error code | Fix |
|---|---|---|
| GPU runs out of memory mid-training | `RUNTIME_GPU_OOM` | Automatic — Backpropagate halves the batch size and retries up to 3 times. To opt out: `Trainer(oom_recovery=False)`. To force smaller: `--batch-size 1`. |
| HuggingFace returns 401 / "model not found" | `DEP_MODEL_LOAD_FAILED` | `huggingface-cli login` and retry. For typos, copy the exact ID from <https://huggingface.co/models>. |
| `register_with_ollama` connection refused | `DEP_OLLAMA_REGISTRATION_FAILED` | Start the daemon: `ollama serve`. Install from <https://ollama.com>. Retryable. |
| Disk full during checkpoint save | `STATE_CHECKPOINT_INVALID` | Atomic writes leave a `.partial` directory on crash — safe to delete. The previous good checkpoint is intact. |
| Training paused on GPU overheat | `RUNTIME_GPU_TEMPERATURE_CRITICAL` | Automatic — Backpropagate pauses on the temperature threshold and resumes as the GPU cools. Improve airflow if it keeps happening. |
| `backprop ui --share` rejected | `RUNTIME_UI_AUTH_NOT_ENFORCED` | Pass `--auth user:password`, or use SSH port-forwarding instead (see [Web UI](#web-ui)). |
| GGUF export failed on first try | `RUNTIME_GGUF_EXPORT_FAILED` | `pip install backpropagate[export]`; on Windows you also need Visual C++ Build Tools + CMake. |

## Reporting bugs

When something fails, Backpropagate prints a line at startup like `run_started run_id=<uuid>` and binds the same ID to every log line, every checkpoint, and every Weights & Biases entry. **Include the `run_id` in any bug report** — it lets a maintainer correlate everything for that exact run.

A good bug report includes:

1. **The `run_id`** — the UUID printed at startup. One UUID lets a maintainer correlate every log line, every checkpoint, and every Weights & Biases entry for that exact run.
2. **The error code** — the `[CODE_NAME]: message` line in stderr. See [error codes](https://mcp-tool-shop-org.github.io/backpropagate/handbook/error-codes/) for the catalog of stable codes.
3. **The redacted traceback.** Stderr is automatically redacted in non-verbose mode (Bearer tokens, `sk-*`, `hf_*`, AWS keys, `password=` / `token=` / `api_key=` pairs are scrubbed) — safe to paste. For the full unredacted traceback, re-run with `BACKPROPAGATE_DEBUG=1` (or `--verbose`); review before posting.
4. **The `backprop info` output.** One command prints Python / PyTorch / CUDA / GPU model / VRAM / OS / installed extras — everything the maintainer needs to bisect a platform-specific regression.

The [bug report template](https://github.com/mcp-tool-shop-org/backpropagate/issues/new?template=bug_report.yml) prompts for each of these explicitly so triage moves fast. Questions, ideas, or "is this expected?" threads belong in [GitHub Discussions](https://github.com/mcp-tool-shop-org/backpropagate/discussions). Security issues should be reported privately via the [GitHub Security Advisory](https://github.com/mcp-tool-shop-org/backpropagate/security/advisories/new) form — see [SECURITY.md](SECURITY.md) for the policy and response timelines.

## Privacy

All training happens locally on your GPU. Backpropagate makes no network requests except to download models from HuggingFace (which you initiate). No telemetry, no cloud dependency.

## References

Backpropagate's defaults and multi-run training mode are built on recent research. If you're interested in the underlying techniques:

- **Hu et al. 2021.** *LoRA: Low-Rank Adaptation of Large Language Models.* [arXiv:2106.09685](https://arxiv.org/abs/2106.09685) — the foundational paper introducing LoRA, which is how Backpropagate trains adapters efficiently.
- **Biderman et al. 2024.** *LoRA Learns Less and Forgets Less.* [arXiv:2405.09673](https://arxiv.org/abs/2405.09673) — empirical evidence that LoRA at rank 256 with all-linear targets matches full fine-tuning quality on most post-training tasks at 67% of the compute. Drives Backpropagate's v1.3 default LoRA configuration.
- **Thinking Machines 2025.** *LoRA Without Regret.* [thinkingmachines.ai/blog/lora](https://thinkingmachines.ai/blog/lora/) — the practical follow-up identifying the 10× learning-rate-vs-full-FT correction needed at high LoRA rank.
- **Kirkpatrick et al. 2017.** *Overcoming catastrophic forgetting in neural networks.* [arXiv:1612.00796](https://arxiv.org/abs/1612.00796) — the original characterization of why neural networks "forget" earlier training when you fine-tune on new data (EWC — Elastic Weight Consolidation).
- **Wang et al. 2023.** *Orthogonal Subspace Learning for Language Model Continual Learning.* [arXiv:2310.14152](https://arxiv.org/abs/2310.14152) — O-LoRA, an earlier approach to using LoRA for continual learning by constraining new adapters to orthogonal subspaces.
- **Yadav et al. 2023.** *TIES-Merging: Resolving Interference When Merging Models.* [arXiv:2306.01708](https://arxiv.org/abs/2306.01708) — a foundational technique for merging multiple fine-tuned models without interference.
- **Qiao & Mahdavi 2025.** *Merge before Forget: A Single LoRA Continual Learning via Continual Merging.* [arXiv:2512.23017](https://arxiv.org/abs/2512.23017) — the specific algorithm Backpropagate's multi-run merger implements. A December 2025 preprint; Backpropagate is the paper's first known downstream adopter.

## License

MIT — see [LICENSE](LICENSE).

---

<p align="center">
  Built by <a href="https://mcp-tool-shop.github.io/">MCP Tool Shop</a>
</p>
