Metadata-Version: 2.5
Name: nab-bayes
Version: 1.2.0
Summary: Neighbourhood Algorithm Bayes — Bayesian analysis via MCMC and Gibbs Sampling on Voronoi tessellations.
Author: Denis Samatov
License: MIT
License-File: LICENSE
Keywords: bayesian,geophysics,mcmc,neighbourhood-algorithm,voronoi
Requires-Python: >=3.10
Requires-Dist: numba>=0.56.0
Requires-Dist: numpy>=1.20.0
Requires-Dist: pandas>=1.5.0
Requires-Dist: scipy>=1.7.0
Requires-Dist: tqdm>=4.60.0
Provides-Extra: advanced
Requires-Dist: jax>=0.4; extra == 'advanced'
Requires-Dist: pyro-ppl>=1.9; extra == 'advanced'
Requires-Dist: torch>=2.0; extra == 'advanced'
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Provides-Extra: hmc
Requires-Dist: jax>=0.4; extra == 'hmc'
Provides-Extra: nf
Requires-Dist: pyro-ppl>=1.9; extra == 'nf'
Requires-Dist: torch>=2.0; extra == 'nf'
Provides-Extra: viz
Requires-Dist: arviz>=0.12.0; extra == 'viz'
Requires-Dist: matplotlib>=3.5.0; extra == 'viz'
Requires-Dist: openpyxl>=3.0.0; extra == 'viz'
Requires-Dist: plotly>=5.0.0; extra == 'viz'
Requires-Dist: seaborn>=0.12.0; extra == 'viz'
Requires-Dist: statsmodels>=0.13.0; extra == 'viz'
Description-Content-Type: text/markdown

# NAB Python

**Neighbourhood Algorithm Bayes (NAB)** — a tool for Bayesian analysis using Markov Chain Monte Carlo (MCMC) and Gibbs Sampling based on Voronoi tessellations.

## At a glance

| | |
|---|---|
| **What** | Bayesian appraisal of parameter ensembles: MCMC with Gibbs sampling over Voronoi cells of sampled models (Neighbourhood Algorithm, Bayesian stage). |
| **Install** | `pip install nab-bayes` (PyPI); Docker image published by CI. |
| **Quality** | 470+ automated tests; CI on every push. |
| **Performance** | Benchmarks and their limits: [performance summary](docs/reports/performance_gpu_plan/performance_summary.md). |
| **Maintainer** | Denis Samatov (Heriot-Watt TPU Center); see the commit history for contributors. |

## 🚀 Quick Start

### 1. Install Dependencies

Install the published package from PyPI:

```bash
pip install nab-bayes
```

For notebook usage with inline Plotly figures, install the visualization extras:

```bash
pip install "nab-bayes[viz]"
```

For a reproducible developer setup from a repository checkout:

#### ⚡ Automated Installation (Recommended for GPU/CUDA/MPS)
To automatically detect your OS/GPU configuration and install the correct accelerated PyTorch build:
```bash
git clone https://github.com/denis-samatov/neighbourhood-algorithm-bayes.git
cd neighbourhood-algorithm-bayes
python install_gpu.py
```
This script handles the heavy lifting of resolving CUDA on Windows/Linux or MPS on Apple Silicon, then installs all optional components.

#### 📦 Custom/Pip Installation (with CUDA 12.1)
```bash
pip install -r requirements-gpu.txt
```

#### 💻 Standard CPU Installation
```bash
pip install -e ".[dev,viz,nf,hmc]"
```
After installation, run tests to verify:
```bash
pytest -q
```


### 2. Minimum Example

Using the test data provided in `data/`:

```bash
nab run data/punq20/Distributions.txt data/punq20/misfit.tsv \
    output/solutionProb.tsv output/sortedProb.tsv \
    --chains 5 --burn-in 10000 --length 50000 --refresh 1000 \
    --plot --name quickstart_run
```

This will run the MCMC sampling using 5 chains and generate an interactive HTML report. See [User Guide](docs/user-guide/usage.md) for full CLI options.

### 3. Python / Notebook API

```python
import nab

result = nab.fit(
    distribution_file="data/punq20/Distributions.txt",
    misfit_file="data/punq20/misfit.tsv",
    run_dir="output/notebook_run",
    chains=5,
    burn_in=10_000,
    chain_length=50_000,
    sampler="gibbs",
    diagnostics=True,
)

result.posterior.head()
result.plot_posterior()
```

For notebook-oriented usage, see [Python API](docs/user-guide/python_api.md).

## 📚 Documentation

Detailed documentation is available in the [`docs/`](docs/) directory:

1. [**User Guide**](docs/user-guide/usage.md) — Installation, data preparation, CLI usage, and output formats.
2. [**End-to-End Pipeline**](docs/user-guide/pipeline.md) — Post-run workflow for adaptive prior narrowing, Top-N posterior analysis, and related diagnostics.
3. [**Mathematical Description**](docs/reference/mathematics_eng.md) — Theory behind NAB, Voronoi approximation, and MCMC sampling.
4. [**Concepts**](docs/getting-started/concepts.md) — High-level explanation and analogies.
5. [**Python API**](docs/user-guide/python_api.md) — `nab.fit(...)`, `nab.load_run(...)`, `nab.analyze(...)`, and notebook workflows.
6. [**Analytical Validation**](docs/user-guide/analytical_validation.md) — 6 synthetic validation experiments against analytical ground truth.
7. [**Architecture**](docs/reference/architecture.md) — Code structure, data flow, and module descriptions.
8. [**Style Guide**](docs/style-guide/style_guide.md) — Rules for contributing text and docstrings.

## 🏗️ Architecture Overview & Data Flow

NAB is designed as a pipeline with clear separation between data loading, evaluation, simulation, and reporting.

**Data Flow Pipeline:**

1. **Input:** The user provides a `Distributions` file defining parameter bounds/types, and a `Misfit` (solutions) file containing previously evaluated models and their errors.
2. **Initialization:** The `BayesianEvaluator` constructs the Voronoi tessellation space across the bounded parameters.
3. **Sampling:** NAB supports five sampler modes: Gibbs (default), adaptive Gibbs, parallel tempering, surrogate-assisted HMC, and Normalizing Flows. Depending on the mode, the algorithm either walks the Voronoi cells directly or fits a smooth approximation to accelerate posterior exploration.
4. **Validation/Diagnostics:** The chains are evaluated for convergence (R-hat, Effective Sample Size).
5. **Output:** Posterior probabilities are calculated, normalized, and saved to `solutionProb.tsv` and `sortedProb.tsv`. A `chain_history` trace is optional.
6. **Visualization:** If `--plot` is enabled, trace plots and posterior distributions are rendered to an interactive HTML report.

## ⚙️ Configuration & CLI

The primary entry point is the `nab` CLI. You can also invoke programmatically via `nab.core.orchestrator.run_analysis`.

For notebook-first usage, prefer the higher-level public API:

```python
import nab

result = nab.fit(...)
analysis = nab.analyze(result, top_n=20)
```

Supported sampler modes are `gibbs`, `adaptive`, `tempered`, `hmc`, and `nf`. The [User Guide](docs/user-guide/usage.md) includes a comparison table describing what each sampler does and when to use it.

`gibbs`, `adaptive`, and `tempered` are the reference discrete NAB samplers. `hmc` and `nf` are intentionally marked as approximate / experimental in manifests, reports, and diagnostics.

## Evidence Tiers

NAB distinguishes **baseline_scientific** (gibbs, adaptive, tempered) and **approximate_exploratory** (hmc, nf) evidence tiers.
For full details on trusted benchmark regeneration, verification, and cross-language equivalence criteria, see [Evidence Tiers](docs/reference/evidence_tiers.md).

For CLI arguments, supported distributions, and input formats, see the [User Guide](docs/user-guide/usage.md). For the post-processing workflow, see the [Pipeline Guide](docs/user-guide/pipeline.md).

## 🎯 Practical Interpretation of the Bayesian Result

In this repository, the Bayesian NAB output is positioned as a **starting approximation** rather than a final answer by itself.

Use it as:
1. A warm start (`MAP` / top-probability models) for building an approximation of the target function.
2. A probabilistic region-of-interest selector (where to spend expensive forward evaluations next).
3. An initialization stage before local refinement, surrogate fitting, or physics-constrained optimization.

If you want a narrower follow-up analysis instead of only the raw NAB outputs, the recommended next step is:
1. Run `nab adaptive-prior` to propose a tighter prior with ESS/PSIS checks.
2. Run `nab dashboard --top-n N` to export and visualize the highest-probability `Top-N` subset under the adaptive prior.

## 🧪 Testing & CI

To ensure reliability, NAB maintains an extensive pytest suite and uses Ruff for code quality.

**Run tests:**

```bash
pytest -q
```

**Run linter and formatter:**

```bash
ruff check .
ruff format .
```

**Benchmark `data/nab_data`:**

A reproducible performance benchmark for the prepared `data/nab_data` ensembles
measures data-loading (cold/warm + peak heap) and discrete Gibbs throughput:

```bash
python scripts/benchmark_nab_data.py --dataset static_maas --burn-in 100 --length 1000
python scripts/benchmark_nab_data.py --dataset dyn --no-sampling          # loading only
```

The script is read-only with respect to `data/` and warms up Numba JIT before
timing. Pass `--json out.json` to capture machine-readable results.


*Note on Java Parity:* This is a Python port of the original Java implementation. Posterior-level equivalence is validated rather than trajectory identity. See [Evidence Tiers](docs/reference/evidence_tiers.md#cross-language-equivalence-java--python) for acceptance criteria.

## 🚑 Troubleshooting

**Top issues:**

1. **`DistributionParseError`:** The `Distributions.txt` format is extremely strict. Ensure no trailing whitespaces on lines and use tab/space delimiters as documented.
2. **NaNs in output:** Check if the misfit values in your input TSV are too large, leading to numerical underflow. Use the `--normalize` flag (enabled by default) to handle large scales.
3. **MCMC Chains taking too long:** Reduce `--length` for a quick test, or increase `OMP_NUM_THREADS` / `NUMBA_NUM_THREADS` env vars to leverage multi-core CPU capabilities in the Voronoi kernels.
4. **No interactive plot:** Ensure you added `--plot` and check your console for the `file://` link generated.

## 🛡️ Security

Do not commit sensitive data or proprietary configurations into `data/` or `output/`. Outputs are added to `.gitignore`, but double-check before pushing any custom project data. No secret keys are required to use this software.

## 🔬 Origins

This repository is a Python port of the **Java implementation of NAB**, which itself is a re-implementation of the original **NA-Bayes** algorithm by **Malcolm Sambridge** (Australian National University).

**Founding papers:**

1. Sambridge, M. *Geophysical Inversion with a Neighbourhood Algorithm I: searching a parameter space.* Geophys. J. Int., **138**, 479–494, 1999.
2. Sambridge, M. *Geophysical Inversion with a Neighbourhood Algorithm II: appraising the ensemble.* Geophys. J. Int., **138**, 727–746, 1999.
