Metadata-Version: 2.4
Name: tmc-pbp
Version: 0.1.0
Summary: Pseudo-Boolean Polynomial decomposition for data analysis
Author-email: Tendai Mapungwana Chikake <tendaichikake@phystech.edu>
License: MIT
Project-URL: Repository, https://github.com/Tenfleques/tmc-pbp
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy
Requires-Dist: pandas
Requires-Dist: scipy
Requires-Dist: bitarray
Requires-Dist: scikit-learn
Provides-Extra: viz
Requires-Dist: matplotlib; extra == "viz"
Provides-Extra: images
Requires-Dist: opencv-python; extra == "images"
Requires-Dist: Pillow; extra == "images"
Provides-Extra: server
Requires-Dist: flask; extra == "server"
Provides-Extra: all
Requires-Dist: matplotlib; extra == "all"
Requires-Dist: opencv-python; extra == "all"
Requires-Dist: Pillow; extra == "all"
Requires-Dist: flask; extra == "all"
Dynamic: license-file

# tmc-pbp

Pseudo-Boolean Polynomial (PBP) decomposition for data analysis. Training-free, deterministic algebraic decomposition of data matrices into multilinear polynomials over binary variables.

## What it does

Given a data matrix (rows = variables, columns = observations), PBP decomposes it into a multilinear polynomial where each term represents a specific combination of variables and its coefficient quantifies the interaction strength. The decomposition is:

- **Training-free** -- no learned parameters, no optimization
- **Deterministic** -- same input always produces the same output
- **Fast** -- sub-millisecond per sample, 15K genes in 1.7 seconds
- **Interpretable** -- each coefficient names a specific variable interaction
- **Mathematically equivalent** to the Walsh-Hadamard spectral transform

## Install

```bash
pip install tmc-pbp            # core (numpy, pandas, scipy, bitarray, scikit-learn)
pip install "tmc-pbp[all]"     # plus matplotlib, OpenCV, Pillow and Flask for the viz, image and server modules
```

The package is imported as `pbp`. A pure-Python core is always available (`pbp.core`).
The optional C backend (`pbp.core_c`) is shipped as source; compile it once in the installed
package directory to enable it:

```bash
bash "$(python -c 'import pbp, os; print(os.path.dirname(pbp.__file__))')/build_pbp.sh"
```

## Quick start

### Python API

```python
from pbp.core import create_pbp, pbp_vector

import numpy as np
matrix = np.array([
    [5.2, 4.8, 3.1, 2.0],  # Gene A
    [3.0, 3.5, 4.2, 4.8],  # Gene B
    [1.0, 1.5, 2.8, 4.5],  # Gene C
])

# Full decomposition -> DataFrame with (y, coeffs, degree)
pbp = create_pbp(matrix)

# Fixed-length vector representation (2^m - 1 elements)
vec = pbp_vector(matrix)
```

### CLI

```bash
pbp analyze matrix.csv --output results/   # Full decomposition + energy profile
pbp vector matrix.csv                       # Fixed-length PBP vector
pbp hasse matrix.csv --format dot           # Hasse diagram (Graphviz DOT)
pbp info matrix.csv                         # Matrix stats
pbp version                                 # Package version
```

### Web application

```bash
bash pbp/webapp/run.sh
# Open http://localhost:8430
```

6-tab interface covering decomposition, epistasis, quorum sensing, sequential inference, structural comparison, and anomaly detection. 25 API endpoints. OpenAPI docs at `/docs`.

## Modules

| Module | Purpose |
|--------|---------|
| `core.py` | Canonical PBP decomposition. `create_pbp()`, `pbp_vector()`, permutation/coefficient/variable matrices. Bitarray-based, no size limit. |
| `inference.py` | PBP-DAG sequential inference. Autoregressive sampling through the Hasse diagram using PBP coefficients as Boltzmann energies. Greedy, stochastic, and beam search modes. |
| `evaluate.py` | Evaluation metrics (Spearman, Kendall, NDE), bootstrap CIs, and baseline predictors (expression magnitude, PCA order, random). |
| `pertpy_integration.py` | Pertpy/scverse-compatible functions: `pbp_distance()`, `pbp_score()`, `pbp_classify_type()`. Works with AnnData objects or plain numpy arrays. |
| `cli.py` | Command-line interface. `pbp analyze\|vector\|hasse\|info\|version`. |
| `pipeline.py` | High-level pipeline utilities. |
| `server.py` | Legacy Flask server for DAG inference visualization. |
| `webapp/` | FastAPI web application (25 endpoints, single-page frontend with Chart.js + D3.js). |

## Applications

### Genetic epistasis (Perturb-seq)
PBP spectral profiles characterize interaction structure in combinatorial perturbation screens. Degree-wise energy separates interaction type from magnitude. Validated on 4 datasets from 4 labs (Norman, Wessels, Dixit, Joung).

```python
from pbp.pertpy_integration import pbp_score, pbp_classify_type
scores = pbp_score(adata, groupby='perturbation', reference='control')
```

### Quorum sensing signal integration
Decompose factorial QS experiments into main effects (a1, a2) and interaction (a12) per gene. Gate classification: AND, OR, antagonistic, mixed.

### Sequential inference
Predict activation order from expression matrices using greedy sampling through the PBP Hasse diagram.

```python
from pbp.inference import pbp_dag_sample
result = pbp_dag_sample(pbp_df, m, mode='greedy')
print(result['sequence'])  # Predicted activation order
```

### Anomaly detection
PBP total energy flags samples with disrupted interaction structure. Interpretable: identifies which specific interactions are affected.

### Structural comparison
Pairwise Hasse diagram distance compares interaction architectures across samples, conditions, or organisms.

## Key functions

| Function | What it does |
|----------|-------------|
| `create_pbp(matrix)` | Full PBP decomposition -> DataFrame with monomial index, coefficient, degree |
| `pbp_vector(matrix)` | Fixed-length vector (2^m - 1 coefficients) for machine learning |
| `pbp_dag_sample(pbp, m)` | Greedy/stochastic/beam sequential inference through Hasse diagram |
| `evaluate_sequence(pred, gt)` | Spearman rho, Kendall tau, top-k accuracy, NDE |
| `pbp_distance(adata, groupby)` | Pairwise structural distance (Pertpy-compatible) |
| `pbp_score(adata, groupby)` | Spectral profile + interaction fraction per group |
| `pbp_classify_type(adata, groupby)` | Synergistic/antagonistic/additive classification |

## Mathematical identity

PBP decomposition is mathematically identical to:
- **Walsh-Hadamard spectral transform** (Fourier analysis on {0,1}^m)
- **Epistatic interaction coefficients** (Poelwijk et al. 2019)
- **Multilinear extension** of pseudo-Boolean functions (Hammer & Rudeanu 1968, Boros & Hammer 2002)

## Citation

```bibtex
@article{chikake2025compoptics,
  author  = {Chikake, Tendai M. and Goldengorin, Boris I. and Pardalos, Panos M.},
  title   = {Pseudo-Boolean Polynomial Approach to Solving Computer Vision Tasks},
  journal = {Computer Optics},
  volume  = {49},
  number  = {6},
  pages   = {1191--1201},
  year    = {2025},
  doi     = {10.18287/COJ1815},
}
```

## License

MIT
