Metadata-Version: 2.1
Name: termux-bitnet
Version: 2.0.1
Summary: 1.58-bit (i2_s) BitNet on-device inference SDK with verified ARM NEON & DotProd kernels for Android Termux & ARM64
Author-email: uno-km <dev@amevafoundation.org>
License: Apache-2.0
Keywords: bitnet,llm,arm64,termux,quantization,neon,dotprod
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: C++
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: Android
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: requests >=2.28.0
Requires-Dist: ameva-runtime >=2.0.0
Requires-Dist: ameva-component-sdk <2.0,>=0.1.0

# Termux-BitNet (v2.0.1 Sovereign Ternary)

> **Production-Grade Universal 1.58-Bit (i2_s) On-Device LLM Inference Engine with ARM64 NEON DotProd SIMD & Native Vulkan GPU Acceleration.**

<p align="center">
  <a href="https://pypi.org/project/termux-bitnet/"><img src="https://img.shields.io/pypi/v/termux-bitnet?color=3775A9&logo=pypi&logoColor=white&label=PyPI" alt="PyPI Version"></a>
  <a href="https://www.npmjs.com/package/termux-bitnet"><img src="https://img.shields.io/npm/v/termux-bitnet?color=CB3837&logo=npm&logoColor=white&label=npm" alt="npm Version"></a>
  <a href="https://github.com/uno-km/termux-bitnet/releases/tag/v2.0.1"><img src="https://img.shields.io/github/v/release/uno-km/termux-bitnet?color=0969da&logo=github&logoColor=white&label=Release" alt="GitHub Release"></a>
  <a href="https://uno-km.vercel.app/lib/bitnet/"><img src="https://img.shields.io/badge/Docs-Official%20Portal-004499.svg?logo=googlechrome&logoColor=white" alt="Documentation Portal"></a>
  <a href="https://opensource.org/licenses/Apache-2.0"><img src="https://img.shields.io/badge/License-Apache_2.0-blue.svg?logo=apache&logoColor=white" alt="License"></a>
  <img src="https://img.shields.io/badge/Core-ARM64%20NEON%20%7C%20DotProd%20%7C%20Vulkan-brightgreen.svg?logo=cplusplus&logoColor=white" alt="Core Acceleration">
  <img src="https://img.shields.io/badge/Models-BitNet%20%7C%20Falcon--E%20%7C%20Falcon3--7B-purple.svg" alt="Supported Models">
  <img src="https://img.shields.io/badge/Fleet-Galaxy%20S25%20%7C%20S20%20%7C%20A53%20%7C%20A35-orange.svg" alt="Tested Hardware">
</p>

---

### 🏛️ Executive Engineering Disclosure: v2.0.1 Generational Breakthroughs

`termux-bitnet v2.0.1` extends the sovereign ternary architecture with **hallucination-prevention prompt wrapping**, **hardware-level GPU watchdog defense**, and **dynamic LM Head memory slicing**, delivering fully calibrated conversational generation across diverse 1.58-bit models.

#### 📚 Official Academic & Upstream Credibility
* **Foundation Technical Whitepaper**: [AOSF-TR-2026-BITNET-TERNARY-02: Root Cause Analysis of Ternary Numerical Collapse and Dynamic Activation Engine](https://uno-km.vercel.app/labs/?menu=research-papers)
* **Upstream Contributions to Microsoft**:
  * [microsoft/BitNet #551](https://github.com/microsoft/BitNet/pull/551): Resolved ARM QK=128 stride memory layout desynchronization and word salad.
  * [microsoft/BitNet #624](https://github.com/microsoft/BitNet/pull/624): Completed 1x4_32W parallel NEON `sdot` hardware acceleration kernel and automated Android Termux tooling.
* **Upstream Contributions to GGML**:
  * [ggml-org/whisper.cpp #4089](https://github.com/ggml-org/whisper.cpp/pull/4089): Mobile heterogeneous GPU-Encoder / CPU-Decoder split-mode architecture.

---

### 📑 Supported Pretrained Models & Optimal Settings

| Model Identifier | Parameter Count | Disk Footprint | Activation Kernel | Recommended Prompt Template | Recommended Acceleration Flags |
|---|:---:|:---:|:---:|:---:|---|
| **Microsoft BitNet 2B-4T** | 2.0B | 1.13 GB | Squared ReLU (`--act-fn relu2`) | `--prompt-template raw` | `-ngl 30` (100% GPU Offload) |
| **TII Falcon-E 1B Instruct** | 1.0B | 635 MB | SwiGLU (`--act-fn swiglu`) | `--prompt-template falcon` | `-ngl 24` (100% GPU Offload) |
| **TII Falcon3 7B Instruct** | 7.45B | 3.05 GB | SwiGLU (`--act-fn swiglu`) | `--prompt-template chatml` | `--chunk-layers 4 --vocab-slice 32768` |
| **BitNet Embedding 270M** | 268M | 367 MB | Linear / RMSNorm | `--prompt-template raw` | CPU Zero-Copy Mmap (30+ tok/s) |

---

### 🛠️ Comprehensive CLI & SDK Parameter Reference Manual

The inference runtime provides user-directed controls to eliminate hallucinations and tailor execution to mobile hardware:

| Option Flag | Type / Range | Default | Purpose & Architectural Behavior |
|---|:---:|:---:|---|
| `--prompt-template` | `chatml` \| `falcon` \| `raw` \| `none` | `none` | Wraps user input in model-compliant dialogue tokens to eliminate drift and word salad. |
| `--act-fn` | `auto` \| `relu2` \| `swiglu` | `auto` | Enforces mathematical activation kernel routing (`relu(x)^2` vs `SiLU(x) * up`). |
| `--chunk-layers` | Integer (0 = disabled, e.g. 4) | `0` | Submits GPU command buffers in chunks of N layers, bypassing the 2.5s Mali watchdog fence. |
| `--vocab-slice` | Integer (0 = disabled, e.g. 32768) | `0` | Slices FP16 LM Head projection rows, slashing VRAM consumption by 576MB on 7B models. |
| `-ngl`, `--n-gpu-layers` | Integer (0 to total layers) | `0` | Number of transformer layers permanently offloaded into Vulkan GPU VRAM. |
| `--prompt-prefix` | String | `""` | Custom prefix prepended to user prompt before tokenization. |
| `--prompt-suffix` | String | `""` | Custom suffix appended to user prompt before tokenization. |
| `--eos-token-id` | Integer | Model default | Explicit override for End-of-Sequence token ID to guarantee generation termination. |

---

### ⚡ Empirical Real-Device Hardware Fleet Scorecard (Ground Truth)

All benchmarks were empirically measured on genuine Samsung Galaxy hardware under unrooted Android Termux Bionic libc environments with steady thermal equilibrium and verified semantic outputs:

| Target Model | Test Device & Hardware | Offload Mode | Token Generation Speed | Verified Semantic Output String | Status |
|---|---|:---:|:---:|---|:---:|
| **BitNet-2B** | **Galaxy S25** (Snapdragon 8 Elite) | GPU (30/30) | **19.37 tok/s** | `"Paris. Cathy has a lot more money..."` | **PASS** |
| **BitNet-2B** | **Galaxy S20** (Turnip Adreno 650) | GPU (30/30) | **7.71 tok/s** | `"Paris. Cathy has a lot more money..."` | **PASS** |
| **BitNet-2B** | **Galaxy A35** (Exynos 1380 Mali-G68) | GPU (30/30) | **4.22 tok/s** | `"Paris. Cathy has a lot more money..."` | **PASS** |
| **BitNet-2B** | **Galaxy A53** (Exynos 1280 Mali-G68) | GPU (30/30) | **3.26 tok/s** | `"Paris. Cathy has a lot more money..."` | **PASS** |
| **Falcon-E-1B** | **Galaxy S25** (Snapdragon 8 Elite) | GPU (24/24) | **34.35 tok/s** | `"a complex and often subject to numerous..."` | **PASS** |
| **Falcon-E-1B** | **Galaxy S20** (Turnip Adreno 650) | GPU (24/24) | **10.76 tok/s** | `"a complex and often subject to numerous..."` | **PASS** |
| **Falcon-E-1B** | **Galaxy A35** (Exynos 1380 Mali-G68) | GPU (24/24) | **5.81 tok/s** | `"a complex and often subject to numerous..."` | **PASS** |
| **Falcon-E-1B** | **Galaxy A53** (Exynos 1280 Mali-G68) | GPU (24/24) | **4.46 tok/s** | `"a complex and often subject to numerous..."` | **PASS** |
| **Falcon3-7B** | **Galaxy S25** (Snapdragon 8 Elite) | GPU (28/28) | **8.30 tok/s** | `"the only power that can be considered..."` | **PASS** |
| **Falcon3-7B** | **Galaxy S20** (Turnip Adreno 650) | GPU (-ngl 8) | **2.62 tok/s** | `"a large city in the north-west region..."` | **PASS** |
| **Falcon3-7B** | **Galaxy A53** (Exynos 1280, 6GB RAM) | GPU Chunked (4) | **1.74 tok/s** | `"the only power that can be considered..."` | **PASS** |
| **Falcon3-7B** | **Galaxy A35** (Exynos 1380, 6GB RAM) | GPU Chunked (4) | **0.78 tok/s** | `"the only power that can be considered..."` | **PASS** |

---

## 📦 Installation

### 1. Python SDK (PyPI)
```bash
# Inside Android Termux (prerequisites: clang cmake python openblas)
pkg update && pkg install -y clang cmake python openblas

# Install Termux-BitNet and AMEVA-Runtime
pip install --upgrade termux-bitnet ameva-runtime
```

### 2. Node.js CLI (npm)
```bash
npm install -g termux-bitnet @ameva/runtime
```

### 3. One-Touch Native Installer
```bash
curl -sSL https://raw.githubusercontent.com/uno-km/termux-bitnet/main/install.sh | bash
```

---

## 🚀 Quickstart Recipes

### 1. Python Programmatic Inference
```python
from termux_bitnet import BitNetEngine, BitNetConfig

# Configure for real-time Falcon-E-1B or BitNet-2B with template wrapping
config = BitNetConfig(
    model_path="~/.cache/termux-bitnet/models/falcon-e-1b-instruct-i2_s.gguf",
    device="gpu",           # "auto", "cpu", or "gpu"
    n_gpu_layers=24,        # 100% GPU offload
    chat_template="falcon", # Formats with <|im_start|>user\n...
    chunk_layers=4,         # Prevent Mali watchdog timeouts
    vocab_slice=32768,      # Slashing LM Head VRAM
    n_threads=4,
    temperature=0.7,
    top_p=0.95
)

with BitNetEngine(config) as engine:
    print("[Prompt]: What is the capital of France?")
    print("[Response]: ", end="", flush=True)
    for token in engine.generate_stream("What is the capital of France?"):
        print(token, end="", flush=True)
    print()
    metrics = engine.get_last_metrics()
    print(f"Speed: {metrics.tokens_per_second:.2f} tok/s")
```

### 2. Standalone CLI Usage
```bash
# 1. Run 1B real-time conversational model with Falcon wrapping
termux-bitnet run -m models/falcon-e-1b-instruct-i2_s.gguf \
  -p "What is the capital of France?" --prompt-template falcon -ngl 24 -t 4

# 2. Run Microsoft BitNet 2B with Squared ReLU
termux-bitnet run -m models/bitnet-2b-ggml-model-i2_s.gguf \
  -p "What is the capital of France?" --act-fn relu2 --prompt-template raw -ngl 30 -t 4

# 3. Run 7.45B model on 6GB RAM device with GPU Chunking and Vocab Slicing
termux-bitnet run -m models/falcon3-7b-instruct-1.58bit-i2_s.gguf \
  -p "Explain quantum computing simply" --prompt-template chatml \
  --chunk-layers 4 --vocab-slice 32768 -ngl 12 -t 4
```

---

## 🌐 Ecosystem Links & Documentation

* **Official Documentation Portal**: [https://uno-km.vercel.app/lib/bitnet/](https://uno-km.vercel.app/lib/bitnet/)
* **Foundation Research Lab**: [https://uno-km.vercel.app/labs/](https://uno-km.vercel.app/labs/)
* **License**: Apache-2.0
* **Governance**: AMEVA Open-Source Foundation (AOSF) & @uno-km
