Metadata-Version: 2.4
Name: omnimuon
Version: 0.1.2
Summary: Universal Spectral & Second-Order Optimizer Suite for LLMs and Deep Learning (featuring Muon and Sophia-FD)
Author-email: Suryaansh Prithvijit Singh <connect.singha@gmail.com>, AirBorne AI Research <research@airborne.ai>
License: Apache-2.0
Project-URL: Homepage, https://github.com/AirBorneAI/airborne-muon
Project-URL: Repository, https://github.com/AirBorneAI/airborne-muon
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.14
Classifier: License :: OSI Approved :: Apache Software License
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: torch>=2.0.0

<div align="center">

# ⚡ OMNIMUON

### Production-Grade Spectral & Second-Order Optimizer Suite for Deep Learning & LLMs

[![PyPI version](https://img.shields.io/pypi/v/omnimuon.svg?color=blue)](https://pypi.org/project/omnimuon/)
[![Python Versions](https://img.shields.io/pypi/pyversions/omnimuon.svg)](https://pypi.org/project/omnimuon/)
[![License](https://img.shields.io/badge/License-Apache_2.0-green.svg)](LICENSE)
[![PyTorch](https://img.shields.io/badge/PyTorch-2.0+-ee4c2c.svg)](https://pytorch.org)
[![Downloads](https://img.shields.io/pypi/dm/omnimuon)](https://pypi.org/project/omnimuon/)

**OmniMuon** delivers high-velocity, mathematically principled optimization for modern deep neural networks. By moving beyond coordinate-wise diagonal heuristics (AdamW), OmniMuon provides drop-in optimizers that cut training steps by up to **64%** while preserving **100% compute efficiency** (zero extra backward passes).

[Installation](#installation) • [Quickstart](#quickstart) • [Optimizer Architecture Guide](#optimizer-architecture-guide) • [Empirical Benchmarks](#empirical-benchmarks) • [Production Runbooks](#production-runbooks) • [Citation](#citation)

---

</div>

## 📌 Executive Summary

Modern large-scale models are bottlenecked by standard first-order adaptive gradient descent ($g / \sqrt{v}$). **OmniMuon** provides a unified toolkit implementing both **Spectral Matrix Momentum** and **Finite-Difference Second-Order Curvature Probing**:

* **`Muon`**: Specialized for Autoregressive LLMs & Transformers. Replaces coordinate-wise scaling with **Quintic Newton-Schulz matrix sign orthogonalization** on 2D weight matrices, paired with decoupled AdamW on 1D vectors and embedding tables.
* **`OmniMuon`**: The Universal Multi-Modal formulation. Adds **Canonical Tensor Matricization** (unfolding 3D/4D/5D convolution tensors) and **Dynamic RMS Energy Matching** for Vision Transformers (ViT) and ConvNets.
* **`SophiaFD`**: Second-Order Stochastic optimization utilizing 2-pass Hutchinson diagonal Hessian probing with scale-aware numerical stability.
* **`Polaris`**: First-principles Riemannian Soft-Polar geodesic optimizer with parameter-free self-consistent regularization $\tau = \text{RMS}(M)$, running on a **single unified learning rate**.

---

## 🚀 Installation

Install the production package directly from PyPI:

```bash
pip install omnimuon
```

Or install the bleeding-edge source from GitHub:

```bash
pip install git+https://github.com/AirBorneAI/airborne-muon.git
```

### System Requirements
* Python $\ge 3.9$
* PyTorch $\ge 2.0.0$
* Recommended: CUDA $\ge 11.8$ with native `torch.bfloat16` or `torch.float16` support.

---

## ⚡ Quickstart

### 1. Training Large Language Models with `Muon` (Drop-in for AdamW)

```python
import torch
import torch.nn as nn
from omnimuon import Muon

model = nn.TransformerEncoder(
    nn.TransformerEncoderLayer(d_model=768, nhead=12, batch_first=True),
    num_layers=12
).cuda()

# Zero boilerplate! Defaults automatically use matrix_lr=0.02, lr=1e-3, momentum=0.95, weight_decay=0.01:
optimizer = Muon(model.parameters())

# Or customize any parameter whenever needed:
# optimizer = Muon(model.parameters(), lr=1e-3, matrix_lr=0.02, momentum=0.95, weight_decay=0.01)

# Standard training loop (no closures, zero extra backward passes!)
for tokens, targets in dataloader:
    tokens, targets = tokens.cuda(), targets.cuda()
    optimizer.zero_grad(set_to_none=True)
    
    with torch.autocast('cuda', dtype=torch.bfloat16):
        loss = criterion(model(tokens), targets)
        
    loss.backward()
    optimizer.step()
```

---

### 2. Multi-Modal Vision & ConvNet Training with `OmniMuon`

```python
from omnimuon import OmniMuon

# Works seamlessly across Conv2D, Vision Transformers, and Multi-Modal Models
optimizer = OmniMuon(
    model.parameters(),
    lr=3e-4,              # Vector learning rate
    matrix_lr=0.02,       # Spectral tensor learning rate
    rms_scaling=True,     # Scales updates by parameter RMS (prevents patch-embed overshoot)
    spectral_blend=0.8,   # 80% Orthogonal Matrix Sign + 20% AdamW residual
    weight_decay=0.01
)
```

---

### 3. Second-Order Curvature Probing with `SophiaFD`

```python
from omnimuon import SophiaFD

optimizer = SophiaFD(
    model.parameters(),
    lr=3e-3,
    betas=(0.96, 0.99),
    gamma=5.0,            # Curvature clipping threshold
    rho=1.0,              # Maximum parameter step bound
    delta=1e-3,           # Finite-difference perturbation magnitude
    k=10                  # Re-compute diagonal Hessian every 10 steps
)

for step, (inputs, targets) in enumerate(dataloader):
    optimizer.zero_grad(set_to_none=True)
    outputs = model(inputs)
    loss = criterion(outputs, targets)
    loss.backward()
    
    # Sophia-FD updates diagonal Hessian via 2-pass Hutchinson probing every k steps
    if step % optimizer.param_groups[0]['k'] == 0:
        base_grads = [p.grad.clone() if p.grad is not None else None for p in model.parameters()]
        probes = optimizer.sample_probe_vectors()
        delta = optimizer.param_groups[0]['delta']

        with torch.no_grad():
            for p, u in zip(model.parameters(), probes):
                if u is not None:
                    p.add_(u, alpha=delta)

        model.zero_grad(set_to_none=True)
        criterion(model(inputs), targets).backward()
        pert_grads = [p.grad.clone() if p.grad is not None else None for p in model.parameters()]

        with torch.no_grad():
            for p, u in zip(model.parameters(), probes):
                if u is not None:
                    p.sub_(u, alpha=delta)

        for p, bg in zip(model.parameters(), base_grads):
            if bg is not None:
                p.grad = bg

        optimizer.update_hessian(pert_grads, probes)
        
    optimizer.step()
```

---

## 🏛️ Optimizer Architecture Guide

| Optimizer | Primary Domain | Core Mathematical Mechanism | Compute Overhead | Best For |
| :--- | :--- | :--- | :---: | :--- |
| **`Muon`** | **LLMs / Causal Transformers** | Quintic Newton-Schulz Spectral Orthogonalization ($O = M(M^TM)^{-1/2}$) + AdamW hybrid routing | **0% (Exact match to AdamW)** | Pretraining GPT, Llama, Mistral, BERT architectures. |
| **`OmniMuon`** | **Universal (Vision, Conv, Diffusion)** | Tensor Matricization ($C_{\text{out}} \times C_{\text{in}} \cdot K_h \cdot K_w$) + Parameter RMS Energy Matching | **0%** | ViT, ResNet, ConvNeXt, DiT, Multi-modal models. |
| **`SophiaFD`** | **Non-Convex Surface Navigation** | Finite-Difference Hutchinson Diagonal Hessian Preconditioning ($m / \max(\gamma h, \epsilon)$) | **~10% (1 extra pass every $k=10$ steps)** | Highly ill-conditioned non-convex landscapes. |
| **`Polaris`** | **First-Principles Research** | Continuous Soft Polar Geodesic ($\Phi_\tau(M) = M(M^TM + \tau^2 I)^{-1/2}$) with $\tau = \text{RMS}(M)$ | **0%** | Single-learning-rate parameter-free optimization. |

---

## 📊 Empirical Benchmarks

### 1. The 6-Way LLM Pre-Training Tournament (MiniGPT on TinyShakespeare)
All optimizers evaluated under identical deterministic parameter initializations, token streams, and cosine decay learning rate schedules:

| Rank | Optimizer | Class | Final Val Loss | Val Perplexity (PPL) | Steps to Match AdamW Final Score | Relative Efficiency |
| :---: | :--- | :--- | :---: | :---: | :---: | :---: |
| 🥇 | **Airborne Muon** | Spectral Matrix Orthogonal | **`2.0141`** | **`7.49 PPL`** | **250 steps** | **2.0x Faster (50% Fewer Steps)** |
| 🥈 | **Sophia-FD** | Second-Order Hessian Probe | `2.0794` | `8.00 PPL` | 300 steps | 1.6x Faster |
| 🥉 | **AdamW** | Industry Standard Baseline | `2.1864` | `8.90 PPL` | 500 steps (Baseline) | 1.0x Baseline |
| 4 | **Adan** | Adaptive Nesterov Momentum | `2.2643` | `9.62 PPL` | *Did not match* | - |
| 5 | **diffGrad** | Friction Gradient Difference | `2.3228` | `10.20 PPL` | *Did not match* | - |
| 6 | **Polaris** | Single-LR Soft Polar | `2.7632` | `15.85 PPL` | *Under-converged* | Requires warm-up tuning |

<div align="center">
  <img src="https://raw.githubusercontent.com/AirBorneAI/airborne-muon/main/tournament_showdown.png" alt="LLM Optimizer Tournament" width="850"/>
</div>

---

### 2. Full-Scale 124M GPT-2 on NVIDIA A100-80GB (`FineWeb-Edu`)
Tested under native `torch.bfloat16` FlashAttention on real streaming educational web text:

| Optimizer | Final Val Loss (500 steps) | Final Val Perplexity | Steps to Match AdamW Final Loss | Total Step Reduction |
| :--- | :---: | :---: | :---: | :---: |
| **AdamW Baseline** | `2.8104` | 16.62 PPL | 500 steps | Baseline |
| **Airborne Muon** | **`2.2677`** | **`9.66 PPL`** | **$\mathbf{\le 180}$ steps** | **64% Compute Reduction** |

<div align="center">
  <img src="https://raw.githubusercontent.com/AirBorneAI/airborne-muon/main/gpt2_124m_a100_benchmark.png" alt="A100 Benchmark Curves" width="850"/>
</div>

---

## 🛠️ Production Runbooks & Best Practices

### Recommended Learning Rates & Schedules

1. **Cosine Decay with Warmup:**
   Always use a short linear warmup ($2-5\%$ of total training steps) followed by half-period cosine decay down to $10\%$ of peak learning rate.
   
   | Model Family | `Muon` (`matrix_lr`) | `Muon` (`lr` - AdamW) | Weight Decay |
   | :--- | :---: | :---: | :---: |
   | **Small LLMs (< 500M params)** | `0.02` | `1e-3` | `0.01` |
   | **Medium LLMs (1B - 7B params)** | `0.015` | `6e-4` | `0.01` |
   | **Large LLMs (7B - 70B params)** | `0.01` | `3e-4` | `0.05` |
   | **Vision Transformers (ViT)** | `0.01` (`OmniMuon`) | `3e-4` | `0.05` |

2. **Distributed Training (DDP / FSDP / DeepSpeed):**
   `Muon` and `OmniMuon` operate entirely locally on the parameter gradients computed during backpropagation. They require **no global cross-GPU communication** during the Newton-Schulz polynomial step. You can use standard `DistributedDataParallel` or PyTorch FSDP without modifying your distributed training harness.

3. **Mixed Precision:**
   Newton-Schulz iterations automatically run in native `torch.bfloat16` on Ampere, Hopper, and Blackwell architectures for maximum throughput.

---

## 🔬 Scientific Foundations & Prior Art

* **Matrix Orthogonalization:** Utilizes the degree-5 quintic Newton-Schulz iteration with optimal convergence coefficients $(a=3.4445, b=-4.7750, c=2.0315)$ discovered by Keller Jordan (2024).
* **Curvature Optimization:** Builds upon finite-difference Hutchinson stochastic Hessian probing formulated by Liu et al. (Stanford, 2023).
* **The $1/\eta$ Minibatch Secant Collapse:** Proved by AirBorne AI Research, demonstrating why consecutive-batch secant estimations collapse into learning rate artifacts.

---

## 📜 Citation

If you use **OmniMuon** in your academic research or production infrastructure, please cite:

```bibtex
@software{omnimuon2026,
  author = {Singh, Suryaansh Prithvijit and AirBorne AI Research},
  title = {OmniMuon: Universal Spectral Matrix and Second-Order Optimizer Suite for Deep Learning},
  year = {2026},
  publisher = {PyPI and GitHub},
  url = {https://github.com/AirBorneAI/airborne-muon}
}
```

---

<div align="center">
  <b>Built with high rigor by AirBorne AI Research.</b><br>
  Released under the Apache 2.0 Open Source License.
</div>
