Metadata-Version: 2.4
Name: nxrt-ep-cuda
Version: 0.1.0.dev6
Summary: ONNX Runtime plugin execution provider (CUDA 13, EXPERIMENTAL) — bundled cdylib for register_execution_provider_library
Author: Justin Chu
License: MIT
Project-URL: repository, https://github.com/justinchuby/onnx-genai
Keywords: onnx,onnxruntime,execution-provider,cuda,plugin
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Rust
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: Microsoft :: Windows
Classifier: Environment :: GPU :: NVIDIA CUDA :: 13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.8
Description-Content-Type: text/markdown
Requires-Dist: nvidia-cuda-runtime==13.1.80
Requires-Dist: nvidia-cublas==13.1.1.3
Requires-Dist: nvidia-cufft==12.1.0.78
Requires-Dist: nvidia-nvjitlink==13.1.115
Requires-Dist: nvidia-cuda-nvrtc==13.1.115
Requires-Dist: nvidia-cuda-cupti==13.1.115
Requires-Dist: nvidia-cudnn-cu13==9.24.0.43
Provides-Extra: onnxruntime
Requires-Dist: onnxruntime>=1.22; extra == "onnxruntime"

# nxrt-ep-cuda (EXPERIMENTAL / PRE-RELEASE)

> ⚠️ **Experimental.** The CUDA execution provider bundled here has been
> validated on physical CUDA hardware (NVIDIA H200, driver 580.105.08,
> CUDA 13.0): it loads and runs the Muse-Glimmer-30B (int4) decoder end-to-end
> with **zero CPU fallbacks**, and the bundled plugin `.so` registers and
> executes ONNX graphs through ONNX Runtime on-device. It remains
> **pre-release**: APIs and packaging may change without notice, and it is not
> yet recommended for production.

A pip-installable **ONNX Runtime plugin execution provider (CUDA 13)**. The
wheel bundles the compiled `onnx-runtime-ep-cuda-plugin` shared library
(`libonnx_runtime_ep_cuda_plugin.{so,dll}`) built with the `cuda` cargo feature
and exposes the absolute path to it so ONNX Runtime can load it via
`RegisterExecutionProviderLibrary(registration_name, library_path)`.

The bundled library exports the ORT plugin-EP C ABI
(`CreateEpFactories` / `ReleaseEpFactory`). It is **not** a Python extension.

## Requirements

- The validated **CUDA 13.1** runtime line. The wheel pins the same NVIDIA
  runtime packages as `requirements-cuda-dev.txt`, including cuBLASLt, cuFFT,
  NVRTC, CUPTI, and cuDNN, so they are installed automatically. It does not
  require `nvcc` or a system CUDA toolkit. The NVIDIA **driver**
  (`libcuda.so.1`) remains a host prerequisite.
- Linux (x86_64) and Windows (AMD64) only. There is no macOS build.

## Install

```bash
pip install nxrt-ep-cuda   # pre-release; may require --pre
```

## Usage

```python
import nxrt_ep_cuda

path = nxrt_ep_cuda.get_library_path()   # absolute path to the bundled cdylib

import onnxruntime as ort
so = ort.SessionOptions()
nxrt_ep_cuda.register(so)                # thin wrapper over register_execution_provider_library
```

`register()` prefers `SessionOptions.register_execution_provider_library` when a
`SessionOptions` is passed, and falls back to the module-level
`onnxruntime.register_execution_provider_library`. If `onnxruntime` is not
installed it raises a clear `ImportError`.

## License

MIT. See the [onnx-genai](https://github.com/justinchuby/onnx-genai) repository.
