Metadata-Version: 2.4
Name: arom-continuum
Version: 1.0.0
Summary: Continuum AI: A hardware-agnostic, zero-copy Deep Learning inference engine bypassing CUDA. Developed under AROM Labs.
Home-page: https://github.com/silenceminds/continuum
Author: Aditya "Aadi"
Maintainer: AROM Labs
Project-URL: Source Code, https://github.com/silenceminds/continuum
Project-URL: Organization, https://github.com/silenceminds
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: C++
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Operating System :: POSIX :: Linux
Requires-Python: >=3.8
Description-Content-Type: text/markdown
Requires-Dist: numpy>=1.21.0
Dynamic: author
Dynamic: classifier
Dynamic: description
Dynamic: description-content-type
Dynamic: home-page
Dynamic: maintainer
Dynamic: project-url
Dynamic: requires-dist
Dynamic: requires-python
Dynamic: summary

# Continuum AI

Developed under AROM Labs by Aditya "Aadi", Continuum AI is a sovereign, framework-agnostic, hardware-accelerated deep learning inference engine. It operates completely independently of massive proprietary toolkits (like PyTorch or the full CUDA stack), executing multi-billion parameter models via dynamic OpenCL linking, direct PTX assembly injection, and extreme quantization.

## Core Architecture

Continuum AI is built on four hardened engineering pillars:

*   **Universal Hardware Abstraction:** Bypasses static linking to proprietary vendor SDKs. Uses dynamic runtime linking (`dlopen`) to probe and lock onto available hardware accelerators (NVIDIA, AMD, Intel) at startup.
*   **Direct Silicon Injection:** Injects raw inline PTX assembly (`mma.sync.aligned.m16n8k8`) directly into the compilation pipeline, forcing NVIDIA architectures (Turing/Ampere) to execute matrix math on physical Tensor Cores without relying on `cuBLAS`.
*   **Zero-Copy Model Ingestion:** Utilizes custom OS-level memory-mapping (`mmap`) to ingest multi-gigabyte `.safetensors` files instantly, streaming weights directly from disk to GPU VRAM with zero intermediate Python allocations.
*   **Extreme Quantization Fabric:** Supports native FP32, FP16, INT8, and W4A16 (packed 4-bit weights) execution pipelines. Dynamic dequantization handles 4-bit and 8-bit packed weights on the fly, slashing VRAM consumption by up to 75%.

## Installation

Install directly from PyPI. The package compiles the C++ OpenCL and PTX backend for your specific local hardware during installation:

```bash
pip install continuum-ai

Requirements: Linux, Python 3.8+, NumPy, and system OpenCL drivers (libOpenCL.so).Hardware Performance DiagnosticsLive benchmark results executing a standard heavy workload ($1024 \times 1024$ Tiled GEMM Matrix Multiplication, 30 consecutive passes) on an NVIDIA Tesla T4:Execution EngineCompute TierLatency (ms/iter)Speedup vs CPUPyTorch (Baseline)CPU (AVX2 / 16 Cores)17.84 ms1.0xContinuum AIFP32 (General ALU)5.00 ms3.5xContinuum AIFP16 (Fast Math)3.93 ms4.5xContinuum AITensor Core (Inline PTX)1.95 ms9.1xContinuum AIW4A16 (4-bit Packed)0.39 ms45.7xContinuum AIINT8 (Quantized Math)0.28 ms63.7xQuick Start: Python APIContinuum provides a Python API that wraps the low-level C++ stateful memory orchestrator.Standard Engine ExecutionPythonimport numpy as np
import continuum_py as cnp

# Initialize the stateful hardware bridge
engine = cnp.Engine()

# Prepare matrices
M, K, N = 1024, 1024, 1024
A_u16 = np.random.randn(M, K).astype(np.float16).view(np.uint16)
B_u16 = np.random.randn(K, N).astype(np.float16).view(np.uint16)
C_u16 = np.zeros((M, N), dtype=np.float16).view(np.uint16)

# Execute via physical Tensor Cores (Sub-millisecond PTX Injection)
engine.gemm_tensor_core(A_u16, B_u16, C_u16)

# Execute in-place neural primitives
engine.gelu(C_u16)
Production Inference Layers (W4A16)Continuum provides high-level neural network wrappers capable of routing data through the extreme quantization fabric.Pythonfrom continuum import InferenceLinear
import continuum_py as cnp
import numpy as np

engine = cnp.Engine()

# Initialize a W4A16 layer (2 INT4 weights packed per byte)
layer = InferenceLinear(engine, in_features=4096, out_features=4096, precision="w4a16")

# Input tensor
x = np.random.randn(1, 4096).astype(np.float16).view(np.uint16)

# Forward pass (Executes the gemm_w4a16 OpenCL kernel)
output = layer(x)
Zero-Copy Safetensors LoaderLoad multi-gigabyte models instantly without RAM bloat:Pythonimport continuum_py as cnp

loader = cnp.SafetensorsLoader()
loader.load("model.safetensors")

# Returns a direct, zero-copy NumPy view over the mmap'd binary buffer
weights = loader.get_tensor("blocks.0.attn.q_proj.weight")
