Metadata-Version: 2.4
Name: umf-format
Version: 0.2.0
Summary: Ultra-Compact Macromolecular Format: Ultra-compact, zero-copy binary container for proteins, nucleic acids, and ligand complexes
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24.0
Requires-Dist: scipy>=1.10.0
Requires-Dist: gemmi>=0.6.0
Requires-Dist: flatbuffers>=23.0
Provides-Extra: gnn
Requires-Dist: torch>=2.0.0; extra == "gnn"
Provides-Extra: test
Requires-Dist: pytest>=7.0.0; extra == "test"
Requires-Dist: biopython>=1.80; extra == "test"
Provides-Extra: benchmark
Requires-Dist: torch>=2.0.0; extra == "benchmark"
Requires-Dist: biopython>=1.80; extra == "benchmark"
Requires-Dist: gemmi>=0.6.0; extra == "benchmark"
Dynamic: license-file

# Ultra-Compact Macromolecular Format (UMF)

> **High-Performance, Zero-Copy Binary Container for Macromolecular Complexes (Proteins, DNA, RNA, Small-Molecule Ligands, & Ions) in Structural Bioinformatics & Deep Learning**

[![C++20](https://img.shields.io/badge/C%2B%2B-20-blue.svg)](https://en.wikipedia.org/wiki/C%2B%2B20)
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
[![FlatBuffers](https://img.shields.io/badge/FlatBuffers-Zero--Copy-green.svg)](https://google.github.io/flatbuffers/)
[![NeRF RMSD](https://img.shields.io/badge/Reconstruction%20RMSD-%E2%89%A4%200.0025%20%C3%85-brightgreen.svg)]()
[![Storage](https://img.shields.io/badge/Compression-98.8%25%20Reduction-orange.svg)]()

---

## 🚀 Overview

The **Ultra-Compact Macromolecular Format (`.umf`)** is a next-generation structural container format designed to replace legacy ASCII formats (`.pdb`, `.cif`) and lossy monomer representations in modern machine learning pipelines, virtual screening, and structural biology archives.

Unlike traditional text-based formats that require expensive string deserialization and runtime distance tree construction, `.umf` leverages **Google FlatBuffers** to enable **direct zero-copy memory mapping** into CPU and GPU PyTorch/NumPy tensors.

### Core Architecture Highlights
1. **Universal Multi-Entity Support**: First-class representation for multi-chain protein complexes, nucleic acids (DNA/RNA), small-molecule drugs/cofactors, metal ions, and crystallographic waters.
2. **Validated Chemical Graph Topology**: Stores explicit covalent bond orders (single, double, triple, aromatic), formal charges, element numbers, and stereochemical parity for small-molecule ligands.
3. **Drift-Free Coordinate Stability ($\le 0.0025\text{ \AA}$)**: Solves the catastrophic kinematic drift of pure dihedral compressors (such as Foldcomp) via **periodic Cartesian anchor snapping** ($\Delta = 15$ residues).
4. **All-Atom NeRF Reconstruction**: Quantized 16-bit rotamer dihedrals ($\chi_{1-4}$) allow deterministic reconstruction of all sidechain heavy atoms within $\le 0.30\text{ \AA}$ RMSD.
5. **Zero-Copy PyTorch Geometric HeteroData Streaming**: Precomputed sparse spatial graphs ($r < 8.0\text{ \AA}$) for intra-protein contacts and bipartite protein-ligand interfaces, eliminating CPU $O(N \log N)$ distance tree generation and cutting dataloading latency by **5.6×**.

---

## 📊 Benchmark Comparisons

### 1. Direct Head-to-Head Comparison (`benchmarks/benchmark_sota.py`)

Evaluated against Foldcomp (`.fcz`), BinaryCIF (`.bcif`), and PDB (`.pdb`) across representative benchmarks:

| Format | Avg Storage (B/res) | Direct PyG Ingestion | Coordinate Drift RMSD | Multi-Chain & Ligands? | Pre-Indexed Graphs? |
|:---|---:|---:|---:|:---:|:---:|
| **Standard PDB (`.pdb`)** | ~1,000 B/res | ~12 ms (text parse + KDTree) | Exact | Text only (no bonds) | ❌ None |
| **BinaryCIF (`.bcif`)** | ~110 B/res | ~15 ms (decompression + parse) | Exact | Tabular | ❌ None |
| **Foldcomp (`.fcz`)** | **~17.9 B/res** | ~12 ms (decompress + KDTree) | **1.96 Å** (up to 17 Å on large proteins) | ❌ Stripped / Crashes | ❌ None |
| **UMF (Backbone Mode)** | **~14.5 B/res** | **~0.35 ms (zero-copy memory map)** | **≤ 0.0025 Å (Drift-Free)** | Monomer fold archive | ❌ None |
| **UMF (All-Atom + Graphs)** | **~35.0 B/res** | **~0.35 ms (zero-copy PyG HeteroData)** | **≤ 0.0025 Å (Backbone) / 0.30 Å (Sidechains)** | ✅ Full RDKit fidelity | ✅ Precomputed $r < 8\text{ \AA}$ |

---

### 2. Downstream Virtual Screening GNN Training (`benchmarks/benchmark_pdbbind_virtual_screening.py`)

5-Epoch training benchmark on 436 co-crystallized complexes from PDBbind:
- **Dataloading I/O Latency**: **5.6× lower** with UMF zero-copy streaming ($1.09\text{ s}$ vs. $6.06\text{ s}$ per epoch).
- **End-to-End Epoch Wall-Clock Time**: **3.3× faster** ($2.14\text{ s}$ vs. $7.13\text{ s}$ per epoch).
- **GPU Dataloader Compute Saturation**: Increases active compute saturation from **15.0%** up to **49.0%** (reducing idle CPU wait from 85% to 51%).

---

## 📦 Installation & Build

### Prerequisites
- CMake $\ge$ 3.20 & Ninja
- C++20 compiler (GCC/G++ $\ge$ 11, Clang $\ge$ 14, or MSVC 2019+)
- Python 3.10+
- Google FlatBuffers (`flatc`)

### Install from PyPI
```bash
pip install umf-format
```

### Building from Source
```bash
# Clone the repository
git clone https://github.com/messiay/Universal_Macromolecular_formate.git
cd Universal_Macromolecular_formate

# Install locally with pip:
pip install -e .
```

---

## 🐍 Python Quickstart

```python
import umf

# 1. Encode any PDB or mmCIF macromolecular complex into .umf binary format
umf.encode_complex("data/complexes/1STP.pdb", "data/complexes/1STP.umf", contact_cutoff=8.0, anchor_interval=15)

# 2. Ingest zero-copy tensors directly into PyTorch Geometric HeteroData
data = umf.to_hetero_tensors("data/complexes/1STP.umf")

# Directly ready for Equivariant GNNs & Virtual Screening
print(data['protein'].coords.shape)                   # (N_prot, 3) Backbone CA coordinates
print(data['protein', 'contact', 'protein'].edge_index) # (2, E_prot) Intra-protein contact graph
print(data['ligand'].coords.shape)                    # (N_lig, 3) Small-molecule coordinates
print(data['ligand', 'bond', 'ligand'].edge_index)    # (2, E_bonds) Covalent chemical bonds
print(data['ligand', 'interacts', 'protein'].edge_index) # (2, E_bip) Bipartite pocket graph
```

---

## 🧪 Testing & Verification

Run the test suite:

```bash
# 1. Test universal macromolecular complex encoding (Biotin-Streptavidin, B-DNA, Hemoglobin)
uv run python -m pytest tests/test_complex_encoding.py -v

# 2. Test adversarial failure cases (disordered loops, negative residues, pure ligands)
uv run python tests/test_adversarial_failure_cases.py

# 3. Test all-atom NeRF reconstruction precision
uv run python tests/test_all_atom_and_features.py

# 4. Run SOTA head-to-head comparison
uv run python benchmarks/benchmark_sota.py 5
```

---

## 📄 License
Licensed under the Apache 2.0 License.
