Metadata-Version: 2.4
Name: foundry-core
Version: 0.1.0rc1
Summary: Foundry: CUDA graph persistence for LLM serving engines (save captured CUDA graphs, restore them at startup)
Project-URL: Homepage, https://github.com/foundry-org/foundry
Project-URL: Source, https://github.com/foundry-org/foundry
Project-URL: Changelog, https://github.com/foundry-org/foundry/blob/main/RELEASE.md
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: C++
Classifier: Operating System :: POSIX :: Linux
Classifier: Environment :: GPU :: NVIDIA CUDA
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Dynamic: license-file
Dynamic: requires-dist

<div align="center">
  <h1>Foundry</h1>
  <p><em>Instantaneous CUDA graph restoration via execution context materialization.</em></p>

  <p>
    <a href="https://www.python.org/"><img alt="Python" src="https://img.shields.io/badge/Python-3.12-blue"></a>
    <a href="https://arxiv.org/abs/2604.06664"><img alt="arXiv" src="https://img.shields.io/badge/arXiv-2604.06664-b31b1b?logo=arxiv&logoColor=white&labelColor=555555"></a>
    <a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/License-Apache_2.0-green.svg"></a>
  </p>
</div>

Foundry is a system that persists CUDA graph states through *template-based context materialization*. It materializes both the structure and execution context of captured CUDA graphs, making graph restoration kernel-agnostic and eliminating the need for hand-crafted patching rules. By intercepting CUDA driver calls, Foundry enforces a **deterministic memory layout** and automatically detects and serializes the **binaries of kernels** used in the CUDA graphs.

With Foundry, LLM serving engines can directly reload CUDA states from disk and skip the warmup process to start in a few seconds.

## Foundry in Action

Demonstration of serving `Qwen/Qwen3-30B-A3B-FP8` with expert parallel size = 2.

<div align="center">
  <img src="assets/foundry_demo.gif" alt="Foundry vs. baseline vLLM cold start" />
  <p><strong>Baseline vLLM (top) vs. vLLM + Foundry (bottom)</strong>. Foundry reconstructs 256 graphs in <strong><u>1 sec</u></strong> (plus ~2 sec sampler warmup + API server init), while original vLLM spends around <strong><u>30 sec</u></strong> to warmup and capture graphs.</p>
</div>

## How it works

Foundry intercepts three classes of CUDA driver calls and routes each through its own piece of in-process state, so SAVE can serialize that state to disk and LOAD can rebuild an identical driver-side state from the archive.

```mermaid
flowchart LR
    subgraph APP["Application calls"]
        direction TB
        A1["cudaMemXXX"]
        A2["cudaModuleLoadXXX<br/>cudaLibraryLoadXXX"]
        A3["cudaGraph capture<br/>cudaGraphLaunch"]
    end

    subgraph FDR["Foundry: indirection &amp; interception"]
        direction TB
        F1["<b>VMM region</b><br/>monotonic cursor →<br/>byte-deterministic offset"]
        F2["<b>Module registry</b><br/>fatbin bytes,<br/>entry_name → CUfunction"]
        F3["<b>Captured graphs</b><br/>each topology group:<br/>1 template + N on-demand"]
    end

    subgraph DRV["CUDA driver / context"]
        direction TB
        D1["cuMemCreate / cuMemMap<br/>cuMemAddressReserve"]
        D2["cuLibraryLoadData<br/>cuLibraryGetKernel"]
        D3["CUgraph<br/>CUgraphExec"]
    end

    A1 --> F1 --> D1
    A2 --> F2 --> D2
    A3 --> F3 --> D3

    classDef app fill:#dfe8ff,stroke:#4060c0,color:#000
    classDef fdr fill:#ffe8d6,stroke:#c08040,color:#000
    classDef drv fill:#d8f0d8,stroke:#40a040,color:#000
    class APP app
    class FDR fdr
    class DRV drv
```

- **Memory.** Every device allocation is funneled into a single VMM region `[base_addr, base_addr + region_size)`. A monotonic cursor gives every tensor a byte-deterministic offset; the underlying physical mapping is done with `cuMemCreate` + `cuMemMap`.
- **Modules / libraries.** As device code is loaded, the fatbin bytes and the `entry_name → CUfunction` table are recorded.
- **Captured graphs.** Each captured `CUgraph` is serialized and then grouped with other graphs that share the same topology (same kernels, dependency DAG and cluster dimensions; they differ only in kernel parameters). One graph per group is kept as the **template**; the rest are stored as **on-demand members** that carry only their per-node kernel parameters.

**SAVE** writes all three pieces to an archive. **LOAD** pre-maps the same VMM range, re-loads the same modules from the packed fatbins, builds each template's `CUgraph` node by node once, and then produces every member by rewriting the template's node parameters, either into a **dedicated** `CUgraphExec` per member (default, eager or lazy) or into one shared exec per template switched with `cuGraphExecUpdate`. Either way restored graphs replay at native speed; kernel handles embedded in the captured graphs resolve to the same device addresses they had at SAVE time. See [`docs/graph-templates.md`](docs/graph-templates.md) for the design and measurements.

## Inference-Engine Integrations

Foundry ships engine integrations under `foundry/python/foundry/integration/`. Per-engine setup instructions live under [`recipe/`](recipe/).

| Engine | Integration code | Documentation | Setup Instructions |
|---|---|---|---|
| SGLang | [`integration/sglang/`](python/foundry/integration/sglang/) | [`docs/sglang/overview.md`](docs/sglang/overview.md) | [`recipe/sglang/README.md`](recipe/sglang/README.md) |
| vLLM | [`integration/vllm/`](python/foundry/integration/vllm/) | [`docs/vllm/overview.md`](docs/vllm/overview.md) | [`recipe/vllm/README.md`](recipe/vllm/README.md) |
| TensorRT-LLM | [`integration/trtllm/`](python/foundry/integration/trtllm/) | [`docs/trtllm/overview.md`](docs/trtllm/overview.md) | [`recipe/trtllm/README.md`](recipe/trtllm/README.md) |

### Status

| Engine | Single GPU | DP | TP | EP |
|---|:---:|:---:|:---:|:---:|
| SGLang | ✅ | ✅ | ✅ | ✅ |
| vLLM | ✅ | ✅ | 🚧 | ✅ |
| TensorRT-LLM | 🚧 | 🚧 | 🚧 | 🚧 |

✅ validated end-to-end (SAVE → LOAD → query) &nbsp;·&nbsp; 🚧 not yet

SGLang TP uses torch symmetric-memory allreduce inside the decode graphs (TP=2/4);
SGLang EP covers DeepEP low-latency (NVSHMEM) and DeepEP v2 (NCCL symmetric
windows). Every SGLang configuration is validated with the full decode-graph set
(batch sizes 1..256) for restore time, per-token latency and greedy-output
equality against unmodified SGLang; see [`recipe/sglang/README.md`](recipe/sglang/README.md#validation).

The adapted SGLang fork is published at [`foundry-org/sglang`](https://github.com/foundry-org/sglang), branch `foundry` (current head `6272eb04c5` = upstream `main` 03ea13a545 + one integration commit; the v0.0.3 pairing `f1d688e52` is kept as `foundry-0.0.3`, the 0.0.2-era integration on `foundry-0.0.2`). The vLLM and TensorRT-LLM forks will follow at `foundry-org/vllm` and `foundry-org/TensorRT-LLM`.

### Performance

SGLang on 8×H100 (2 GPUs per run), all 256 decode graphs captured or restored,
prefill graphs off on both sides, median TPOT of restored vs unmodified SGLang:

| Config | Model | Capture | Restore | Init to `/health`: SGLang → foundry LOAD | TPOT delta (bs 1 / 8 / 32 / 128) |
|---|---|---:|---:|---:|---|
| TP=2 (symm-mem) | Qwen3-32B | 26.7 s | 3.5 s | 63 s → 41 s | -0.4 / -0.3 / +0.0 / -0.7 % |
| EP=2 (DeepEP LL) | Qwen3-30B-A3B | 38.0 s | 2.5 s | 75 s → 45 s | -0.0 / -0.0 / +0.1 / +0.6 % |
| EP=2 (DeepEP v2, NCCL) | Qwen3-30B-A3B-FP8 | 56 s | 2.1 s | 93 s → 45 s | +0.0 / +0.0 / +0.2 / +0.2 % |

Restored graphs keep their programmatic-dependent-launch edges, so per-token
latency matches the captured graphs within run-to-run noise
([`docs/pdl-edge-batching.md`](docs/pdl-edge-batching.md)). Greedy completions
are identical to unmodified SGLang for dense TP and within SGLang's own
run-to-run nondeterminism for MoE.

## Roadmap

See [ROADMAP.md](ROADMAP.md) for the full development plan and progress.

## Requirements

- Linux x86_64, NVIDIA driver with CUDA 12.0+ (the driver's `libcuda.so.1` is loaded at runtime)
- PyTorch: `foundry.ops` is a torch C++ extension bound to the torch it was built
  with (major.minor and CUDA major), like sglang-kernel or flashinfer. Importing
  it under another torch raises a readable error.

## Installation

Prebuilt wheels are published to PyPI as **`foundry-core`** (the import name
is `foundry`). One torch/CUDA pairing per release line:

| foundry-core | torch | CUDA | CPython | Platform |
|---|---|---|---|---|
| 0.1.x | 2.13 (`torch==2.13.*`, cu130 build) | 13.0 | 3.10-3.13 | manylinux_2_28 x86_64 |

```bash
pip install "torch==2.13.0" --index-url https://download.pytorch.org/whl/cu130
pip install "foundry-core>=0.1.0rc1,<0.2"
python -c "import foundry; print(foundry.__version__)"
```

The wheel ships `foundry/ops.*.so` and `foundry/libcuda_hook.so` (the hook
SGLang/vLLM preload) and registers the SGLang plugin entry point. No Boost or
other C++ runtime dependency is needed. Wheels for other torch/CUDA pairs, when
built, are attached to the [GitHub Release](https://github.com/foundry-org/foundry/releases)
with a local version (`0.1.0+cu128.torch2.12`) and install by URL.

### From source

Needed for any other torch, or for development. Requirements:

- CMake 4.0+ and ninja (`pip install "cmake>=4.0" ninja` if the system ones are older)
- the torch you will run with, already installed (Foundry compiles against it; rebuild after changing torch)
- CUDA Toolkit with `nvcc` (CUDA 12+; CUDA 13 with torch cu130)
- Boost headers, header-only (nothing is linked): the vendored copy in
  `third_party/boost` (populated by `tools/release/vendor_boost.sh`), or a
  system Boost >= 1.83 (Ubuntu 24.04: `apt-get install libboost-dev`; conda:
  `conda install -c conda-forge boost-cpp`; or point `FOUNDRY_BOOST_INCLUDE_DIR`
  at the directory that contains `boost/version.hpp`)

```bash
pip install "cmake>=4.0" ninja
# Torch 2.13 with CUDA 13.0
pip install torch==2.13.0 --index-url https://download.pytorch.org/whl/cu130
pip install -e . --no-build-isolation
```

`--no-build-isolation` matters: an isolated build compiles against whatever
torch pip resolves, not the one in your environment. From PyPI's sdist into
an existing environment: `pip install --no-build-isolation --no-binary foundry-core foundry-core`.

Release process (vendoring Boost, tagging, trusted publishing): [`docs/release.md`](docs/release.md).

### Debugging

Verbose C++ logging (allocator hook and graph replay) is gated behind a
single compile-time flag, `FOUNDRY_DEBUG`. Uncomment the
`// #define FOUNDRY_DEBUG` line at the top of `csrc/hook.cpp` or
`csrc/CUDAGraph.cpp` (or build with `-DFOUNDRY_DEBUG`) and reinstall.
With the flag on, every LOAD also logs a per-graph check that the restored
edge records (PDL programmatic edges) survived insertion; see
[`docs/pdl-edge-batching.md`](docs/pdl-edge-batching.md).

## Quick Start

### Graph Capture and Save

Foundry requires LD_PRELOAD to intercept CUDA driver calls. The graph capture and save must run in a subprocess with the hook library preloaded.

```python
import foundry as fdry
import torch

torch.cuda.init()
device = torch.device('cuda:0')
torch.set_default_device(device)

# Set up VMM allocation region for deterministic memory addresses
BASE_ADDR = 0x400000000000
region_size = fdry.parse_size('1GB')
fdry.set_allocation_region(BASE_ADDR, region_size)

# Allocate input tensors
input_a = torch.full((100, 100), 2.0, device=device)
input_b = torch.full((100, 100), 3.0, device=device)

# Warm up the model
model = MyModel()
model(input_a, input_b)
torch.cuda.synchronize()

# Capture CUDA graph
graph = fdry.CUDAGraph()
with fdry.graph(graph):
    result = model(input_a, input_b)

# Replay and verify
graph.replay()
torch.cuda.synchronize()

# Save graph with output tensors
# Produces BOTH graph.json and graph.cugraph (optimized binary format)
graph.save('graph.json', output_tensors=result)

fdry.stop_allocation_region()
```

### Graph Load and Replay

Loading a saved graph also requires LD_PRELOAD and must use the same allocation region base address.

```python
import foundry as fdry
import torch

torch.cuda.init()
device = torch.device('cuda:0')
torch.set_default_device(device)

# Load CUDA modules and libraries from archive
fdry.load_cuda_modules_and_libraries('hook_archive')

# Set up the same allocation region as capture
BASE_ADDR = 0x400000000000
region_size = fdry.parse_size('1GB')
fdry.set_allocation_region(BASE_ADDR, region_size)

# Allocate input tensors (can have different values)
input_a = torch.full((100, 100), 5.0, device=device)
input_b = torch.full((100, 100), 3.0, device=device)

# Load and replay the graph (NOTE: auto-loads .cugraph binary when available)
graph, output_tensor = fdry.CUDAGraph.load('graph.json')
graph.replay()
torch.cuda.synchronize()

# output_tensor now contains the result
fdry.stop_allocation_region()
```

### Async Graph Loading

Load graphs asynchronously with background template building. Graphs with the same topology share a single `CUgraphExec` template — only node parameters are updated before each launch (on-demand replay). Two finish APIs are available:

- `finish_graph_loads(pending)` — bulk, waits for all templates then returns the full list. Simplest call shape.
- `finish_one_graph_load(pending, index)` — per-graph; first call waits on background build completion, later calls just finalize. Use this when you want to interleave finalization with other work (the foundry vLLM integration uses it to walk the same VMM cursor trajectory on LOAD that SAVE recorded).

```python
import foundry as fdry

# Phase 1: parse .cugraph binaries + build topology groups + templates in background
pending = fdry.CUDAGraph.start_graph_builds(
    ["graph_0.json", "graph_1.json", ...], num_threads=24
)

# Background threads now race against whatever the caller does next.
# In practice we found that overlapping with model weight loading is
# net-negative (driver contention), so the vLLM integration kicks off
# start_graph_builds *after* weight load and lets it overlap with the
# cheaper post-load init phases instead.

# Bulk finish:
results = fdry.CUDAGraph.finish_graph_loads(pending)
for graph, output in results:
    graph.replay()

# OR per-graph finish (interleavable):
# for i in range(num_graphs):
#     graph, output = fdry.CUDAGraph.finish_one_graph_load(pending, i)
#     graph.replay()
```

### Graph Manifest and Topology Groups

After capturing all graphs, call `save_graph_manifest()` to group graphs by topology and assign templates. On-demand (non-template) graphs strip dependencies to reduce file size.

```python
import foundry as fdry

# After all graphs are captured and saved
fdry.save_graph_manifest('hook_archive')
```

### Memory Preallocation for Fast Graph Reload

The preallocation API physically allocates memory upfront, enabling subsequent allocations to use a fast path (pointer bump only, no VMM driver calls).

```python
import foundry as fdry

# With preallocation - allocations within 8GB use fast path
with fdry.allocation_region(0x500000000000, '16GB', prealloc_size='8GB'):
    graph, outputs = fdry.CUDAGraph.load('model.json')
    graph.replay()
```

| Function | Description |
|----------|-------------|
| `set_allocation_region(base, size)` | Set VMM allocation region for deterministic memory addresses |
| `stop_allocation_region()` | Stop the allocation region |
| `resume_allocation_region()` | Re-enable a previously stopped allocation region |
| `allocation_region(base, size, prealloc_size=None)` | Context manager to set up VMM allocation region with optional preallocation |
| `preallocate_region(size)` | Manually preallocate memory inside an allocation region |
| `free_preallocated_region()` | Free manually preallocated memory |
| `get_current_alloc_offset()` / `set_current_alloc_offset(offset)` | Read or fast-forward the in-region cursor |
| `parse_size(size)` | Parse a size string (`"1GB"`, `"16MB"`, …) to bytes |
| `load_cuda_modules_and_libraries(archive_dir)` | Load CUDA modules and libraries for graph loading |
| `save_graph_manifest(archive_dir)` | Write graph_manifest.json with topology groups and template assignments |
| `CUDAGraph.start_graph_builds(paths, num_threads)` | Kick off background template build for a list of saved graphs |
| `CUDAGraph.finish_graph_loads(pending)` | Wait for all builds and return `[(graph, output), ...]` |
| `CUDAGraph.finish_one_graph_load(pending, i)` | Finalize one graph by index; interleavable with other work |
| `init_nvshmem_for_loaded_modules()` | After `prepare_communication_buffer_for_model` has bootstrapped NVSHMEM, finalize `nvshmemx_cumodule_init` for each module queued by `load_cuda_modules_and_libraries` |

## Testing

Run the test suite:

```bash
pytest tests/
```

## Setting up clangd

```
conda install -c conda-forge libstdcxx-ng libgcc-ng
conda install -c conda-forge bear
bear -- python setup.py build_ext --inplace
```

## Contributors

- Xueshen Liu
- Yongji Wu
