Metadata-Version: 2.5
Name: cmb-format
Version: 0.3.0
Summary: Specification and reference implementation of CMB (Cell Model Binary), a file format for cell-based meshes and the models defined on them
Project-URL: Homepage, https://github.com/dwfmarchant/cmb-format
Project-URL: Repository, https://github.com/dwfmarchant/cmb-format
Project-URL: Changelog, https://github.com/dwfmarchant/cmb-format/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/dwfmarchant/cmb-format/issues
Project-URL: Specification, https://github.com/dwfmarchant/cmb-format/blob/main/docs/binary-format.md
Project-URL: FormatChangelog, https://github.com/dwfmarchant/cmb-format/blob/main/FORMAT_CHANGELOG.md
Author-email: David Marchant <dwfmarchant@gmail.com>
License-Expression: MIT
License-File: LICENSE
Keywords: binary-format,cmb,file-format,geophysics,mesh,octree
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering
Requires-Python: >=3.11
Requires-Dist: numpy>=1.26
Description-Content-Type: text/markdown

# cmb-format

**CMB (Cell Model Binary)** is a binary file format and Python I/O library
for storing UBC GIF–style tensor and octree meshes and their associated
models. It provides an alternative to the ASCII mesh and model files used
by UBC GIF software. CMB supports uniform and variable-spacing tensor
meshes, octree meshes, multiple named models, and file- and model-level
metadata.

CMB stores mesh geometry and model values as typed arrays, with metadata in
JSON, so consumers can avoid parsing millions of numbers from text. Large
octree consumers that need only arrays can also avoid allocating a full
consumer mesh; see the [discretize round trips and benchmarks](https://github.com/dwfmarchant/cmb-format/blob/main/docs/discretize.md).

## Capabilities

- Store mesh geometry and per-cell model arrays together in a `.cmb` file.
- Store models separately from geometry to avoid duplicating large meshes.
- Read selected model payloads without loading unrequested models.
- Inspect model metadata and file contents without loading model payloads.
- Verify each array's integrity with a SHA-256 checksum.

## Installation

Requires Python 3.11 or newer. NumPy is the only runtime dependency.

Install from PyPI:

```bash
python -m pip install cmb-format
```

Or install from a local checkout:

```bash
python -m pip install .
```

## Usage

The API accepts dictionaries of NumPy arrays describing meshes and models.
CMB cell ordering differs from UBC GIF's, and these routines never reorder
arrays. Tensor meshes number cells x fastest. Embedded octree cells must be
stored in root-local Morton order, which the writers and `read_file` enforce;
see [Octree cell order](#octree-cell-order) and
[Cell numbering / ordering](https://github.com/dwfmarchant/cmb-format/blob/main/docs/binary-format.md#cell-numbering-ordering)
in the format specification.

Write a four-cell tensor mesh and a resistivity model, then read them back:

```python
import numpy as np

import cmb_format as cmb

mesh = {
    "mode": "embedded",
    "mesh_class": "TensorMesh",
    "arrays": {
        "origin": np.zeros(3),
        "h_x": np.array([1.0, 2.0]),
        "h_y": np.array([1.0, 1.0]),
        "h_z": np.array([3.0]),
    },
}
models = {
    "rho": {
        "metadata": {"units": "ohm-m"},
        "array": np.array([10.0, 20.0, 30.0, 40.0]),
    }
}
cmb.write_file("example.cmb", mesh, models)

mesh, models, metadata = cmb.read_file("example.cmb")
rho = models["rho"]["array"]

# Load geometry and only the requested model. File order is preserved.
mesh, models, metadata = cmb.read_file("example.cmb", models=["rho"])
```

`read_file` loads and checksum-verifies all geometry and selected model arrays,
including nested base-mesh geometry. Its three results match `write_file`'s
`mesh`, `models`, and `metadata` parameters, so passing them straight back
preserves the mesh geometry, model arrays, and metadata. `default_padding`
accepts a mapping with `west`, `east`, `south`, `north`, `bottom`, and `top`
keys; omitted keys default to zero. Recognized padding on tensor and uniform
meshes, bare references, and `base_mesh` descriptors is returned as a fresh
complete dictionary of Python integers; an explicit `null` there is omitted.
Unrecognized descriptor fields pass through unchanged on reads and are ignored
by writers. The NumPy arrays are read-only; use `.copy()` if you need to modify
them. `models=None` loads every model, `models=[]` loads none, and duplicate
selections collapse in stored file order.

For inexpensive inspection, use the raw header summaries:

```python
model_summaries = cmb.list_models("example.cmb")
contents = cmb.read_contents("example.cmb")
```

`list_models` returns mappings such as
`{"rho": {"metadata": {"units": "ohm-m"}, "dtype": "float64", "shape": [4]}}`
and reads no array payloads. `read_contents` returns
`{"has_mesh": bool, "mesh_type": str | None, "has_base_mesh": bool,
"n_cells": int, "models": dict}`; it reads only an embedded uniform mesh's
three-value shape array to compute `n_cells`. Nested base-mesh and model
payloads remain unread. Both helpers preserve model metadata, dtype, shape, and
stored order.

To read individual arrays without loading the whole file:

```python
with open("example.cmb", "rb") as f:
    header, data_start = cmb.read_header(f)
    metadata = header["metadata"]
    geometry = cmb.read_arrays(f, data_start, header["mesh"]["arrays"])
    rho = cmb.read_array(f, data_start, header["models"]["rho"]["array"])
```

`read_header` accepts the keyword `read_shape_payload`. Its default `True`
performs full header validation, including the embedded uniform mesh shape
checksum. Set it to `False` for structural header checks without reading shape
payloads; shape values and dependent uniform padding or model-count checks are
deferred. The returned header contains normalized named dictionaries for
recognized padding, including when it reads a legacy v1 file, and uses
`format_version` 2. Unrecognized fields remain unchanged.

### Octree cell order

Embedded octree cells must be stored in root-local Morton order. The base grid
is divided into cubic roots of `L = min(nx, ny, nz)` base cells per axis;
roots are visited x fastest, then y, then z, and the cells in each root follow
the Morton order of their lower corners. `octree_order_keys(position, shape)`
returns each cell's ordering key, and the stored keys must strictly increase.
`write_file` and `build_file_bytes` reject embedded octree geometry that is out
of order or repeats a position; `write_file` does so before opening the output
file. `read_file` rejects such files after verifying the geometry and before
reading any model payload. Nothing reorders arrays, and keeping each model
value aligned with the cell it describes is the caller's responsibility. To
put cells in order, apply one permutation to `level`, `position`, and every
model:

```python
order = np.argsort(cmb.octree_order_keys(position, shape), kind="stable")
```

Reference-mode files store no octree geometry, so their model order cannot be
checked. Store reference models in the root-local order of the mesh they
accompany; matching `n_cells` or `base_mesh` does not show that the cells
match.

`read_header`, `read_array`, and `read_arrays` return stored arrays without
checking cell order, and `list_models` and `read_contents` do not read octree
geometry, so a successful inspection does not certify the order. The low-level
readers can migrate a file written in another order:

```python
with open("unordered.cmb", "rb") as f:
    header, data_start = cmb.read_header(f)
    mesh = header["mesh"]
    mesh["arrays"] = cmb.read_arrays(f, data_start, mesh["arrays"])
    base = mesh["base_mesh"]
    base["arrays"] = cmb.read_arrays(f, data_start, base["arrays"])
    models = {
        name: {**entry, "array": cmb.read_array(f, data_start, entry["array"])}
        for name, entry in header.get("models", {}).items()
    }

keys = cmb.octree_order_keys(mesh["arrays"]["position"], base["arrays"]["shape"])
order = np.argsort(keys, kind="stable")
mesh["arrays"] = {name: values[order] for name, values in mesh["arrays"].items()}
for entry in models.values():
    entry["array"] = entry["array"][order]
cmb.write_file("ordered.cmb", mesh, models, header.get("metadata", {}))
```

Sorting keeps each model value with its cell only if the original file already
paired them correctly. It cannot repair models that are misaligned with the
geometry.

For measured large-octree and tensor round trips and timing methodology, see
[the discretize interoperability notes](https://github.com/dwfmarchant/cmb-format/blob/main/docs/discretize.md). On the measured
2.18-million-cell sample, the generated CMB file is 10.4 MiB versus 28.8 MiB
for UBC, and conversion plus CMB writing was about 21× faster in measurements
taken before octree cell-order validation was added.

## Development

From a local checkout, with pip 25.1 or newer:

```bash
python -m pip install --group dev -e .
python -m pytest
python -m ruff check .
python -m ruff format --check .
```

Committed v1 and v2 reference files in `tests/goldens/v1/` and
`tests/goldens/v2/` test compatibility with the binary format alongside
round-trip tests.

The [format specification](https://github.com/dwfmarchant/cmb-format/blob/main/docs/binary-format.md) defines the file layout
and mesh schemas. Package and format versions are independent; see
[versioning](https://github.com/dwfmarchant/cmb-format/blob/main/docs/binary-format.md#versioning) and the
[package changelog](https://github.com/dwfmarchant/cmb-format/blob/main/CHANGELOG.md)
and [format changelog](https://github.com/dwfmarchant/cmb-format/blob/main/FORMAT_CHANGELOG.md).
