Metadata-Version: 2.4
Name: g2n
Version: 1.14.0
Summary: Optimize and run PyTorch models: an open-core compiler (fusion, buffer planning, persistent compile cache), a quantum circuit simulator (g2n.quantum), plus a license-gated serving platform that runs your models behind an inference server.
Author: g2n
Maintainer: g2n
License: Apache-2.0
Project-URL: Homepage, https://g2n.dev
Project-URL: Documentation, https://g2n.dev/docs
Project-URL: Source, https://github.com/nactttch/g2n
Project-URL: Issues, https://github.com/nactttch/g2n/issues
Project-URL: Benchmarks, https://g2n.dev/benchmarks
Project-URL: Pricing, https://g2n.dev/pricing
Keywords: pytorch,compiler,triton,gpu,inference,serving,model-server,torch.compile,quantum,quantum-simulator,statevector
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Compilers
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: cryptography>=41.0
Provides-Extra: torch
Requires-Dist: torch>=2.4; extra == "torch"
Provides-Extra: triton
Requires-Dist: triton>=2.2; extra == "triton"
Provides-Extra: native
Requires-Dist: g2n-native<0.2,>=0.1; (sys_platform == "linux" and platform_machine == "x86_64") and extra == "native"
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: twine>=5; extra == "dev"
Requires-Dist: pytest>=7; extra == "dev"

# g2n — run models that don't fit, on the GPU you already have

A PyTorch compiler (custom FX fusion passes, a Triton LayerNorm kernel, a
persistent compile cache) plus a license-gated runtime that serves models and
runs local LLMs **larger than your VRAM**. Also ships `g2n.quantum`, a
statevector quantum circuit **simulator** (classical simulation — not quantum
hardware).

Measured on a single RTX 4050 Laptop (6 GB) with
[`benchmarks/bench.py`](https://github.com/nactttch/g2n/blob/main/benchmarks/bench.py)
in this repo — run it on your own card and post what you get:

- **EuroLLM-9B at int4 weights, generating at 2.89 tok/s on a 6 GB card.**
  33 of its 42 layers stay resident; the rest stream from pinned host RAM with
  the transfers overlapped against compute. The planner sizes that split
  automatically and lands within **0.3%** of the best split found by sweeping
  every option by hand.
- **3.5× faster cold compile** than stock `torch.compile` on a post-reboot run.

```bash
pip install g2n torch
```

```python
import torch, g2n

compiled = g2n.compile(model)     # == torch.compile(model, backend="g2n")
y = compiled(x)
```

Free (Community) gives you the g2n fusion passes on stock Inductor and
quantum circuits up to 24 qubits. A license key unlocks more — activation is
one command and then fully offline:

```bash
export G2N_LICENSE_KEY=G2N-XXXX-XXXX-XXXX
g2n activate && g2n status
```

| Unlock | Tier |
|---|---|
| Persistent compile cache (warmup once per machine, not per run) | Pro |
| Enhanced planner: epilogue fusion + custom Triton kernels | Pro |
| Serving platform (`pip install g2n-enterprise`): registry, HTTP node, quantization incl. weight-only int8, CUDA graphs | Pro |
| Quantum: unlimited qubits + circuit fusion | Pro |
| max-autotune, dynamic batching, batched quantum sweeps, model zoo | Enterprise |

## Run a `.g2n` packaged model — free, no license

A `.g2n` pack is a model that is already quantized on disk, so opening it is an
mmap instead of a re-quantization. **Reading one is free and unlicensed** and
lives right here in the Apache-2.0 package:

```python
from g2n.pack import load_packed, inspect

print(inspect("qwen3-8b-int4.g2n")["precision"])   # works with no torch installed
model, manifest = load_packed("qwen3-8b-int4.g2n")
```

The loader builds the module tree on `device="meta"` (allocating nothing), then
binds the packed tensors straight from the memory-mapped file — no fp16
intermediate and no quantization work at load. Creating a pack (`g2n pack`) is a
paid feature; running one never is.

## Quantum in 20 seconds

```python
import g2n.quantum as qf
c = qf.Circuit(2).h(0).cnot(0, 1)      # Bell state
c.measure_all(shots=1000)              # {'00': ~500, '11': ~500}
c.expectation("ZZ")                    # tensor(1.)
```

## Guarantees

- **Never worse than eager**: any compile failure returns your unmodified
  model with a one-line warning.
- **Offline after activation**: license tokens are Ed25519-verified locally;
  no phone-home during runs.
- **Honest numbers**: no benchmark claim without named hardware and a
  script to reproduce it — including on your own machine.

Docs, pricing, benchmarks: **[g2n.dev](https://g2n.dev)** · full manual at
**[g2n.dev/docs](https://g2n.dev/docs)** · questions: **support@g2n.dev**.

## What's open and what isn't

This repository **is** the free core: the compiler, the fusion passes and the
quantum simulator, Apache-2.0, no obfuscation. The LLM offload and serving
numbers quoted above come from `g2n-enterprise`, which is proprietary and paid.
Nothing in either package phones home during a run.
