Metadata-Version: 2.4
Name: voltabf16
Version: 0.1.0
Summary: Working bfloat16 on NVIDIA V100: exact BF16 GEMM on Volta FP16 tensor cores
Author: Misekai LLC
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/misekaii/Convolt
Project-URL: Documentation, https://convolt.readthedocs.io
Project-URL: Source, https://github.com/misekaii/Convolt
Keywords: bfloat16,v100,volta,sm_70,tensor-cores,cuda,pytorch,gemm
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: GPU :: NVIDIA CUDA
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: torch>=2.0
Requires-Dist: nvidia-cutlass<4,>=3.5
Provides-Extra: bench
Requires-Dist: torchvision; extra == "bench"
Provides-Extra: docs
Requires-Dist: sphinx>=7; extra == "docs"
Requires-Dist: furo; extra == "docs"
Requires-Dist: myst-parser; extra == "docs"
Provides-Extra: test
Requires-Dist: pytest; extra == "test"
Dynamic: license-file

# Exact BF16 GEMM on FP16 tensor cores (V100 / sm_70)

CUDA port + extensions of "Exact BF16 Product Packing into FP16 Matrix
Multiplication". Volta has FP16 tensor cores but no BF16 support; this suite
provides BF16 GEMM as a menu of speed/accuracy modes, all built on the same
CUTLASS-sm70 mainloop with an empirically probed accumulator mapping.

| Mode (impl header)        | passes | speed @9216x4096x4096 | accuracy |
|---------------------------|--------|-----------------------|----------|
| scaled_impl (scaled FP16) | 1      | 5.6 ms, 6x FP32 route | == FP32 route on real ML data; survives extreme exponents |
| sigcut_impl (exact 1:1)   | 1      | 5.4 ms                | bit-exact (block-separable exponents, full BF16 value set) |
| gexp_impl (general exp.)  | 1-9 adaptive | 5.4-55 ms       | bit-exact to 16-bit/row exponent spread; 50x better than FP32 route |

- `src/` implementation: `common.cuh` (helpers), `exact_mma.cuh` (CUTLASS config,
  rebase mainloop, probe), `rowscan.cuh`, one `*_impl.cuh` per mode.
- `tests/` correctness: `make test` runs bit-exact/accuracy validation of all
  three modes (zeros, subnormals, Inf/NaN, odd shapes, permutation path).
- `experiments/` benchmarks (`bench_*.cu`) and historical prototypes:
  `bf16_pack_gemm.cu` (paper's 9:8 K<=32 repro), `deep_k_pack.cu`,
  `full_k_pack.cu` (12:11 limbs), `sig_gemm.cu` (WMMA 1:1),
  `cublas_roofline.cu`, `cutlass_hgemm_ref.cu` (attribution baselines).
- `analysis/` real-ML BF16 studies (SmolLM2 weights/activations).
- `python/` the `voltabf16` PyTorch bindings (see below), `docs/` their
  Sphinx site.

Build on the GPU host: `make -j4 && make test` (needs CUDA 12.6 for sm_70 and
a CUTLASS checkout; see `CUTLASS`/`NVCC` vars).

## PyTorch bindings: `voltabf16`

`pip install voltabf16` exposes the `scaled_impl` mode to PyTorch, so bfloat16
training on a V100 uses the FP16 tensor cores instead of torch's non-tensor-core
fallback:

```python
import voltabf16

with voltabf16.autobf16():
    loss = model(x.bfloat16()).sum()
    loss.backward()
```

Nothing in the model changes -- a `TorchDispatchMode` intercepts bfloat16
`aten::mm`/`aten::addmm` below autograd, so forward and both backward GEMMs are
routed. MNIST MLP (784 -> 4096x3 -> 10), batch 2048, 5 epochs:

| mode | ms/step | vs FP32 | test acc |
|---|---|---|---|
| FP32 | 39.2 | 1.00x | 97.93% |
| FP16 | 8.4 | 4.68x | 97.90% |
| BF16 (torch native) | 51.6 | **0.76x** | 97.93% |
| BF16 (voltabf16) | 11.9 | **3.30x** | 97.91% |

4.3x faster than torch's bfloat16, and FP32-class accurate (~1e-7 relative
error against float64, versus ~1.7e-3 native) because the products are exact.
It does not reach FP16 parity and cannot: the mainloop is ~78-83% of cuBLAS
FP16, and the exponent scan and pack passes are inherent. `docs/performance.md`
has the full accounting.

The binding is built against this repository's own `src/*.cuh` -- there is one
copy of those headers, shared with `make test`. `setup.py` copies them into
wheels at build time so an installed wheel is self-contained; an editable
install or a run from the checkout uses `src/` directly.

```bash
pip install -e '.[test,docs,bench]'
pytest                                    # needs a V100; skips otherwise
make -C docs html
python python/examples/bench_mnist.py --epochs 5 --batch 2048 --lr 0.15
make bin/tile_sweep && ./bin/tile_sweep 2048     # CUTLASS tile-shape sweep
```

`.github/workflows/docs.yml` builds and publishes the docs;
`.github/workflows/pypi.yml` publishes `voltabf16` to PyPI via Trusted
Publishing when a GitHub Release is cut.
