Metadata-Version: 2.1
Name: windowsml-llama-core
Version: 2.7.2021a0
Summary: llama.cpp backend core for the Windows ML Runtime
Author: Microsoft Corporation
License: Copyright (C) Microsoft Corporation. All rights reserved.
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Requires-Python: >=3.9
Requires-Dist: windowsml==2.7.2021a0
Description-Content-Type: text/markdown

# Windows ML llama.cpp backend for Python

These distributions add the optional llama.cpp backend to the `windowsml` Python package.
With them installed, the Windows ML Runtime loads GGUF models.

| Distribution | Contents |
|---|---|
| `windowsml-llama` | Metapackage. Installs the core and the CUDA backend. |
| `windowsml-llama-core` | The llama.cpp adapter, the llama.cpp and ggml modules, and the CPU backend. |
| `windowsml-llama-cuda` | The NVIDIA CUDA backend. Depends on NVIDIA's `nvidia-cublas` package. |

Install the metapackage through the `windowsml` extra:

```powershell
pip install "windowsml[llama]"
```

Or install it directly:

```powershell
pip install windowsml-llama
```

Every distribution pins the exact `windowsml` version it was built with.
The payload is installed into `windowsml/lib`, beside the Runtime it was built for.

The CUDA backend requires an NVIDIA GPU and driver.
It depends on NVIDIA's `nvidia-cublas` wheel at the exact version Windows ML validates.
That wheel is published by NVIDIA under the NVIDIA software license, which installing it accepts.
Windows ML loads that cuBLAS, or a later 13.x release, before the backend.
Without a supported GPU, the CUDA backend is skipped and the CPU backend remains available.
