Metadata-Version: 2.4
Name: loggetta
Version: 0.5.0
Summary: Plan and run Mixture-of-Experts fine-tuning on one GPU: check the machine, fit the model, and keep a report of each run.
Author: Jordan Anderson
License-Expression: MIT
Project-URL: Homepage, https://cerinamroth.com/ml/
Project-URL: Documentation, https://github.com/pjordanandrsn/loggetta#readme
Project-URL: Source, https://github.com/pjordanandrsn/loggetta
Project-URL: Issues, https://github.com/pjordanandrsn/loggetta/issues
Project-URL: Changelog, https://github.com/pjordanandrsn/loggetta/blob/main/CHANGELOG.md
Project-URL: Release Notes, https://github.com/pjordanandrsn/loggetta/releases/latest
Project-URL: Research (Hugging Face), https://huggingface.co/spaces/pjordanandrsn/research
Project-URL: Runtime: experts4bit-qlora, https://pypi.org/project/experts4bit-qlora/
Project-URL: Kernels: grouped-nf4-gemm, https://pypi.org/project/grouped-nf4-gemm/
Keywords: machine-learning,llm,gpu,planning,qlora,moe,mixture-of-experts,fine-tuning,quantization,nf4,consumer-gpu,memory-planning
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: experts4bit-qlora[fast,train]>=0.49.0
Requires-Dist: grouped-nf4-gemm>=0.42.0
Requires-Dist: datasets
Requires-Dist: peft>=0.21.2
Provides-Extra: experts4bit
Provides-Extra: test
Requires-Dist: pytest>=7; extra == "test"
Dynamic: license-file

# Loggetta

**Plan the run. Train the model. Keep the evidence.**

Loggetta checks your hardware, chooses a supported configuration for a Mixture-of-Experts (MoE) model, and runs QLoRA
fine-tuning. It saves the plan and a JSON run report with memory use, timing, and checks that the selected
optimizations actually ran. One install includes the runtime
([experts4bit-qlora](https://pypi.org/project/experts4bit-qlora/)) and the GPU kernels
([grouped-nf4-gemm](https://pypi.org/project/grouped-nf4-gemm/)).

> **Limits.** Plans are estimates, not an out-of-memory guarantee. The released training path is single-GPU MoE
> training on Linux with a supported NVIDIA CUDA GPU. Dense models are planned but not yet supported for training.
> [Dense training](https://github.com/pjordanandrsn/loggetta/blob/main/docs/DENSE.md)

```bash
pip install loggetta
loggetta inspect
loggetta plan Qwen/Qwen3-30B-A3B --seq 2048
```

The plan shows where weights will live, estimated memory use, and why alternatives were rejected. It checks the budget
before downloading model weights. Estimates can miss; they are not an out-of-memory guarantee.

## New in 0.5.0

- **Training plans no longer borrow another model's reserve.** A model with no receipts on this GPU is priced at the
  card's worst measured training slack. In sample, no training plan now sits under its measured peak.
- **Plans record the backend switches their estimate read** (with experts4bit-qlora 0.52.0 or later), and
  `loggetta execute` warns when the running process differs.
- **Dense plans:** an estimate no longer falls when one of its terms rises.

Full list: [CHANGELOG](https://github.com/pjordanandrsn/loggetta/blob/main/CHANGELOG.md).

## Train on your data

```bash
loggetta train Qwen/Qwen3-30B-A3B \
  --dataset ./data/train.jsonl --format text \
  --seq 512 --micro-batch 1 --steps 20 --seed 42 \
  --out runs/my-training --adapter-out adapters/my-adapter
```

Local JSONL, JSON, CSV, Parquet and TXT files, Hub datasets, Alpaca instructions and text-only chats are supported.
Data is validated and tokenized before weights load. Chat data trains only the assistant turns by default. Reload the
adapter with `loggetta.load_adapter("adapters/my-adapter")`. See the
[training guide](https://github.com/pjordanandrsn/loggetta/blob/main/docs/TRAINING.md).

To run a saved decision later: `loggetta plan MODEL --out plan.json`, then `loggetta execute plan.json --out runs/`.
Pass earlier run reports back with `--observations runs/` and later plans use those measurements.

## Measured results

The included runtime and kernels do the compute; these are matched training runs, not planner benchmarks.

| Workload | Result |
| :--- | :--- |
| **Qwen3-30B-A3B QLoRA · RTX 5090** | Unsloth spends **1.92×** e4b's GPU time per step, and **2.80×** its wall-clock time on an AMD EPYC 7713 host. Comparable held-out loss; Unsloth peaked lower (24.27 vs 26.16 GB). [Result](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/bench/h2h-2026-10-02/tc1/RESULTS-tc1-pos69.md) |
| **Why two numbers** | GPU time doesn't depend on the host. Unsloth runs about 14× e4b's CPU operations per step, so its wall-clock time grows on a slower host. Earlier wall-clock readings, before e4b's current defaults: 2.352× and 2.468×. |
| **Planner memory check · Qwen3-30B-A3B · RTX 5090** | **24.54 GiB** estimated process peak, **24.34 GiB** measured, after calibration from earlier runs. One in-sample case, not a guarantee. In the MoE audit re-run with training plans no longer borrowing another model's reserve (`evidence/2026-10-09-moe-plan-vs-driver-no-borrow`), today's planner puts no RTX A2000 training plan under its measured peak, in sample (the replan section; older recorded plans were). [Plan vs run](https://github.com/pjordanandrsn/loggetta/blob/main/docs/RESULTS.md) · [MoE audit](https://github.com/pjordanandrsn/loggetta/blob/main/evidence/2026-10-09-moe-plan-vs-driver-no-borrow/README.md) |

The comparison used torch 2.12.1+cu130 and transformers 5.5.0 for both frameworks, with matched adapters, initialization and
tokens. Loggetta does not predict throughput.

## Models

| Model | Tested in Loggetta |
| :--- | :--- |
| OLMoE-1B-7B-0924, Granite-3.1-3B-A800M | training and serving run |
| Granite-4.0-H-tiny | training run; serving refused (Mamba state) |
| Qwen3-30B-A3B, Qwen3.6-35B-A3B, ERNIE-4.5-21B-A3B | planned; serving validated |
| LFM2-8B-A1B | planned; serving refused (conv state) |
| Mixtral-8x7B-Instruct | planned |
| Gemma-4-26B-A4B-it, Nemotron-3.5-Lightning-30B-A3B | supported by the runtime; not yet in Loggetta's sweep |

Every row is supported by the included runtime's QLoRA path. Hybrid models (Qwen3.6, Granite-4.0-H, LFM2, Nemotron-H)
need `--packing concat` with chat or Alpaca data. The runtime's
[capability register](https://github.com/pjordanandrsn/experts4bit-qlora/blob/main/docs/capabilities.json) is the
authority.

## Scope

Single-GPU MoE planning and QLoRA training; serving placement is planned, not launched. Not yet: multi-GPU,
throughput prediction. GPU runs need Linux, an NVIDIA CUDA GPU and a compatible PyTorch. Pre-1.0.

Dense models are not supported for training yet: their plans are estimates that missed a held-out check (a further
reading is pending), and their training runs only behind `--allow-development-executor` until a 24 GB capacity reading.

[GitHub](https://github.com/pjordanandrsn/loggetta) ·
[Results](https://github.com/pjordanandrsn/loggetta/blob/main/docs/RESULTS.md) ·
[Architecture](https://github.com/pjordanandrsn/loggetta/blob/main/docs/ARCHITECTURE.md) ·
[Research and releases](https://cerinamroth.com/ml/)
