Metadata-Version: 2.4
Name: ameva-cluster
Version: 1.0.0
Summary: Symmetric Disaggregated Mobile RAM Pooling & On-Device AI Acceleration Runtime
Author-email: Eunho Kim <contact@uno-km.com>
License: Apache-2.0
Keywords: disaggregated-memory,memory-pooling,on-device-ai,mobile-cluster,vulkan-compute,termux,llama-cpp-rpc,tensor-sharding,fault-tolerance,distributed-inference,guardband-memory,hybrid-streaming
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Dynamic: license-file

# AMEVA-Cluster

[![PyPI](https://img.shields.io/pypi/v/ameva-cluster.svg?style=flat-square&color=0369a1)](https://pypi.org/project/ameva-cluster/)
[![Python](https://img.shields.io/pypi/pyversions/ameva-cluster.svg?style=flat-square)](https://pypi.org/project/ameva-cluster/)
[![npm](https://img.shields.io/npm/v/%40ameva%2Fcluster.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/@ameva/cluster)
[![npm downloads](https://img.shields.io/npm/dm/%40ameva%2Fcluster.svg?style=flat-square&color=b91c1c)](https://www.npmjs.com/package/@ameva/cluster)
[![License](https://img.shields.io/badge/License-Apache_2.0-004499.svg?style=flat-square)](https://github.com/uno-km/ameva-cluster)

> **Symmetric Disaggregated Mobile RAM Pooling & On-Device AI Acceleration Runtime with Compute-Memory Decoupling**

---

## 1. Executive Architecture & Threat Model

Executing frontier AI models (7B–70B parameters, 4B Flow Matching Diffusion, Multimodal Vision) on edge mobile smartphones inevitably triggers immediate **Android Low Memory Killer (LMK) `SIGKILL` terminations** and extreme battery thermal throttling.

Conventional distributed inference frameworks force every connected node to execute computational passes, causing lower-tier edge devices to overheat, desynchronize, and crash.

**AMEVA-Cluster** resolves this architectural limitation through **Compute-Memory Decoupling**:
- **Symmetric Single-Binary Architecture**: Both `master` and `worker` roles run from a unified binary (`ameva-cluster`), dynamically configurable via CLI flags.
- **Pure RAM Pooling**: Worker nodes execute **zero GPU/NPU compute passes**, serving exclusively as remote high-bandwidth LPDDR memory pools to eliminate worker thermal load.
- **Android LMK Safety Guard-Band (300MB)**: Automatically deducts safety buffers from active RAM metrics (`MemAvailable`) to prevent Android kernel out-of-memory kills.
- **Protocol Guard Proxy & MasterTunnel**: Intercepts all incoming connections with a 15-byte magic preamble and 60-second time-windowed HMAC-SHA256 mutual authentication. Drops unauthorized RPC probes in `< 0.001s` (`HTTP 403 Forbidden`).
- **Storage-Assisted Hybrid Failover**: Automatically recovers from worker dropouts by streaming remaining tensor weights directly from local high-speed UFS 3.1 flash memory via kernel `mmap`.

```
[ Master Node (Flagship GPU) ]
    │
    ├── Local Compute: Flagship Adreno / Mali GPU (100% of forward passes)
    │
    ├── MasterTunnel (Loopback: 127.0.0.1:5105X)
    │     └── Transparently injects HMAC-SHA256 handshake token
    │
    ├── USB 3.0 ADB / Tailscale WireGuard Transport (Encrypted)
    │     │
    │     ├──▶ [ Worker 1 (Galaxy S20 - 12GB LPDDR5) ] ──▶ ClusterGuardProxy ──▶ Loopback rpc-server
    │     ├──▶ [ Worker 2 (Galaxy A35 - 8GB LPDDR4X) ] ──▶ ClusterGuardProxy ──▶ Loopback rpc-server
    │     └──▶ [ Worker 3 (Galaxy A53 - 6GB LPDDR4X) ] ──▶ ClusterGuardProxy ──▶ Loopback rpc-server
    │
    └── Storage Failover: Local UFS 3.1 mmap fallback upon node disconnect
```

---

## 2. Installation & Quickstart

### Python (PyPI)
```bash
pip install ameva-cluster
```

### Node.js / TypeScript (npm)
```bash
npm install -g @ameva/cluster
```

### Termux One-Liner (Android ARM64)
```bash
pkg update && pkg install python nodejs-lts git -y
pip install ameva-cluster
```

---

## 3. Command Line Interface (CLI) Reference

### 1) Starting a Worker Daemon (Edge Phone)
Launch the worker daemon with automatic Termux WakeLock acquisition and Protocol Guard protection:
```bash
# Protected worker on public port 50052 (isolates rpc-server on loopback 50055)
ameva-cluster worker --port 50052 --guard-band 300
```
Key Flags:
- `--port <int>`: External listening port (default: `50052`).
- `--isolated-port <int>`: Internal isolated raw RPC loopback port (default: `50055`).
- `--guard-band <int>`: Memory safety guardband in MB (default: `300`).
- `--no-guard`: Disables HMAC authentication proxy (raw debugging only).

### 2) Running the Master Orchestrator (Workstation / Host Phone)
Coordinate distributed inference across multiple edge smartphones:
```bash
# Auto tensor sharding across Galaxy A35 and Galaxy A53
ameva-cluster master \
  -m models/Qwen2.5-7B-Instruct-Q4_K_M.gguf \
  --workers 100.106.251.21:50052,100.77.47.37:50052 \
  -ts auto \
  -p "Explain the physics of gravitational lensing."
```
Key Flags:
- `-m, --model <path>`: Path to target model (`.gguf`, `.safetensors`, `.bin`).
- `--workers <csv>`: Comma-separated list of worker endpoints (`ip:port`).
- `-ts, --tensor-split <ratio>`: Sharding ratio across master and workers (e.g. `50,50` or `auto`).
- `--engine <name>`: Backend engine: `llamacpp`, `diffusion`, `vision`, `stt`, `tts`, `bitnet`, `train`.
- `-p, --prompt <text>`: Inference prompt.

### 3) Pre-Flight Diagnostics & Fleet Probing
Inspect network round-trip time (RTT), packet jitter, and compute node availability:
```bash
ameva-cluster probe --fleet 100.106.251.21:50052,100.77.47.37:50052
```

---

## 4. 6-Modality Distributed Integration

AMEVA-Cluster natively interconnects with the six official on-device AI modality runtimes:

### 1) Large Language Models: `termux-llamacpp`
```bash
termux-llama run -m Qwen2.5-7B-Instruct-Q4_K_M.gguf \
  --rpc 100.106.251.21:50052,100.77.47.37:50052 \
  --tensor-split auto \
  -p "Synthesize distributed systems theory."
```

### 2) Image Diffusion: `termux-diffusion`
```bash
termux-diffusion generate \
  --model z_image_turbo-q4.gguf \
  --vae taef1.gguf \
  --prompt "Hyper-detailed cybernetic macro lens photograph" \
  --rpc 100.77.47.37:50052 \
  --tensor-split 50,50 \
  --steps 8
```

### 3) Multimodal Vision: `termux-vision`
```bash
termux-vision analyze \
  --model Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf \
  --mmproj mmproj-Qwen2.5-VL-7B-Instruct-f16.gguf \
  --image telemetry.png \
  --rpc 100.106.251.21:50052 \
  --prompt "Identify all memory leaks in this timeline graph."
```

### 4) Speech-to-Text: `termux-stt`
```bash
termux-stt transcribe conference_call.wav \
  --model ggml-large-v3-turbo.bin \
  --rpc 100.106.251.21:50052 \
  --tensor-split 50,50
```

### 5) Text-to-Speech: `termux-tts`
```bash
termux-tts synthesize "Distributed edge clustering operational." \
  --voice en_us_neutral \
  --rpc 100.77.47.37:50052 \
  --output alert.wav
```

### 6) 1-Bit LLM & LoRA Training: `termux-bitnet` & `termux-train`
```bash
# BitNet 1.58-bit ternary inference
termux-bitnet run -m bitnet_b1_58-3B.gguf --rpc 100.106.251.21:50052

# Distributed LoRA fine-tuning
termux-train lora --base-model Qwen2.5-3B.gguf --dataset train.jsonl --rpc 100.106.251.21:50052
```

---

## 5. Software Development Kit (SDK) API

### Python API
```python
from ameva_cluster import (
    AmevaCluster,
    ClusterMaster,
    ClusterWorker,
    ClusterHardwareProbe,
    ClusterSecurityVerifier,
    ClusterConfig,
)

# 1. Probe local hardware and subtract 300MB safety buffer
mem = ClusterHardwareProbe.get_memory_info(guard_band_mb=300)
print(f"Usable RAM: {mem['usable_mb']}MB / Total: {mem['total_mb']}MB")

# 2. Acquire Termux system WakeLock
ClusterHardwareProbe.ensure_wakelock()

# 3. Execute disaggregated cluster inference
AmevaCluster.run_master(
    model_path="Qwen2.5-7B-Instruct-Q4_K_M.gguf",
    workers=["100.106.251.21:50052", "100.77.47.37:50052"],
    tensor_split="40,30,30",
    prompt="Explain quantum entanglement."
)
```

### Node.js / TypeScript API
```javascript
const { AmevaCluster } = require('@ameva/cluster');

// 1. Launch protected worker daemon
const worker = AmevaCluster.startWorker({
  port: 50052,
  guardBandMb: 300,
  enableGuard: true
});

// 2. Launch master orchestrator
const master = AmevaCluster.runMaster({
  modelPath: "Qwen2.5-7B-Instruct-Q4_K_M.gguf",
  workers: ["100.106.251.21:50052", "100.77.47.37:50052"],
  tensorSplit: "auto",
  prompt: "Demonstrate distributed Node.js tensor coordination."
});

master.stdout.on('data', (chunk) => process.stdout.write(chunk));
```

---

## 6. Security Gating & License Verification (Fail-Fast E403)

Multi-device distributed clustering is an exclusive capability governed by the AMEVA-Cluster protocol.

When standalone execution is attempted without official licensing or package presence, the runtime halts with code **`E403_CLUSTER_LICENSE_REQUIRED`**:
```
================================================================================
[AMEVA-CLUSTER] CLUSTER LICENSE REQUIRED (E403)
================================================================================
Multi-device distributed clustering is an exclusive capability of 'ameva-cluster'.
Standalone distributed execution without the official AMEVA-Cluster package is prohibited.
================================================================================
```

### Community Tier Inclusion
The official package distribution comes pre-configured with the default community authorization key (`AMEVA-COMMUNITY-CLUSTER-AUTH-TOKEN-2026`). Developers can immediately link consumer phones without manual credential injection, while external unauthorized clients are blocked at the network barrier.

---

## 7. Official Documentation & Portal

- **Web Documentation**: [https://uno-km.vercel.app/lib/cluster/](https://uno-km.vercel.app/lib/cluster/)
- **Guard Protocol Specification**: [https://uno-km.vercel.app/lib/cluster/guard-protocol.html](https://uno-km.vercel.app/lib/cluster/guard-protocol.html)
- **6-Modality Sharding Guide**: [https://uno-km.vercel.app/lib/cluster/modalities-guide.html](https://uno-km.vercel.app/lib/cluster/modalities-guide.html)
- **Fleet Topology & Tuning**: [https://uno-km.vercel.app/lib/cluster/fleet-optimization.html](https://uno-km.vercel.app/lib/cluster/fleet-optimization.html)

---

## 8. License

Licensed under the Apache-2.0 License. Copyright (c) 2026 Eunho Kim ([@uno-km](https://github.com/uno-km)).
All rights reserved.
