Metadata-Version: 2.4
Name: disscube
Version: 0.4.0
Summary: Declarative spatial data cubes: describe sources, grid and derived variables in TOML; get a cataloged, reproducible cube.
Author: Sérgio Souza Costa
Maintainer: Sérgio Souza Costa
License: MIT
Project-URL: Homepage, https://github.com/DisSModel/disscube
Project-URL: Repository, https://github.com/DisSModel/disscube
Project-URL: Documentation, https://dissmodel.github.io/disscube/
Project-URL: Issues, https://github.com/DisSModel/disscube/issues
Project-URL: Changelog, https://github.com/DisSModel/disscube/blob/main/CHANGELOG.md
Keywords: spatial data cube,geospatial,land use and land cover change,brazil data cube,reproducibility
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Topic :: Scientific/Engineering :: GIS
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.0
Requires-Dist: xarray
Requires-Dist: zarr
Requires-Dist: rasterio
Requires-Dist: geopandas
Requires-Dist: shapely
Requires-Dist: scipy
Requires-Dist: numpy
Requires-Dist: pandas
Requires-Dist: pyproj
Requires-Dist: fsspec
Requires-Dist: toml
Requires-Dist: rioxarray
Requires-Dist: affine
Requires-Dist: pooch>=1.8.0
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: pytest-cov; extra == "dev"
Requires-Dist: mypy>=1.13; extra == "dev"
Requires-Dist: ruff<0.17,>=0.16; extra == "dev"
Requires-Dist: ipykernel; extra == "dev"
Provides-Extra: bdc
Requires-Dist: fiona; extra == "bdc"
Requires-Dist: pystac-client; extra == "bdc"
Provides-Extra: s3
Requires-Dist: s3fs; extra == "s3"
Provides-Extra: netcdf
Requires-Dist: h5netcdf; extra == "netcdf"
Requires-Dist: h5py; extra == "netcdf"
Provides-Extra: dissmodel
Requires-Dist: dissmodel<0.7.0,>=0.6.0; extra == "dissmodel"
Provides-Extra: docs
Requires-Dist: mkdocs<2,>=1.6; extra == "docs"
Requires-Dist: mkdocs-material>=9.5; extra == "docs"
Dynamic: license-file

# DisSCube

[![CI](https://github.com/DisSModel/disscube/actions/workflows/ci.yml/badge.svg)](https://github.com/DisSModel/disscube/actions/workflows/ci.yml)
[![Docs](https://github.com/DisSModel/disscube/actions/workflows/docs.yml/badge.svg)](https://dissmodel.github.io/disscube/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

> **Status: Alpha — stable APIs for the core pipeline; declarative models still evolving.**

DisSCube is the spatial data cube engine of the **DisSModel** ecosystem. It converts raw geospatial sources (rasters, vectors) into derived variables aligned to LUCC (Land Use and Cover Change) modeling grids, ready for Cellular Automata models and spatio-temporal analysis.

📖 **Documentation:** <https://dissmodel.github.io/disscube/>

## Core concept

```
SpatialSource  →  Derivation  →  Variable  →  DerivedVariable (Zarr)
```

A **source** (`SpatialSource`) goes through a **derivation** (`SpatialDerivation` or `Derivation`) that applies an **operator** on a **grid** (`GridSpec`), producing a **derived variable** registered in the SQLite catalog and stored in Zarr.

## Installation

```bash
git clone https://github.com/DisSModel/disscube.git
cd disscube
python -m venv .venv && source .venv/bin/activate
pip install -e .
```

## Basic usage

### 1. Initialize the catalog and register a grid

```python
from disscube.client import CubeClient
from disscube.utils.grids import register_local_grid

cube = CubeClient(catalog="catalog.db", store="./data/")

grid = register_local_grid(
    cube,
    name="AC",
    bbox_geo=(-73.99, -11.15, -66.62, -7.11),
    resolution=5_000.0,
)
```

### 2. Register a source

```python
from disscube.models import SpatialSource

cube.register_spatial_source(SpatialSource(
    id="mapbiomas_2020",
    name="MapBiomas Acre 2020",
    format="raster",
    asset_url="data/raw/mapbiomas_2020.tif",
    crs="EPSG:4326",
    time=2020,
))
```

### 3. Derive — declarative mode (recommended)

```python
from disscube.derivation import Derivation

d = Derivation(
    target="forest_pct",
    source_id="mapbiomas_2020",
    operator="percentage",
    class_code=3,
    role="driver",
    valid_from="2020",
    valid_until="2020",
)

cube.derive_declarative(d, grid_id="AC/5km")
```

### 4. Derive — direct mode

```python
from disscube.models import SpatialDerivation, Variable

cube.derive(SpatialDerivation(
    source_id="mapbiomas_2020",
    grid_id="AC/5km",
    role="driver",
    variables=[Variable(name="forest_pct", operator="percentage", class_code=3)],
    valid_from="2020",
    valid_until="2020",
))
```

### 5. Load the result

```python
da = cube.load("forest_pct", grid_id="AC/5km")
print(da.shape)   # (rows, cols)
```

### 6. Get the cube out

```python
ds = cube.to_dataset(["forest_pct", "dist_roads"], grid_id="AC/5km", period=("2015", "2020"))
# xarray.Dataset: (y, x) static and (time, y, x) temporal variables, CRS and transform via ds.rio

cube.export_geotiff(["forest_pct"], "forest.tif", grid_id="AC/5km")   # one band per variable and year
cube.export_netcdf(["forest_pct"], "cube.nc", grid_id="AC/5km")       # CF-1.8; pip install "disscube[netcdf]"
```

Exports carry their provenance: each GeoTIFF band and netCDF variable records the
`spec_hash`, the `content_hash` of the stored data and the `source_checksum` of the
input it came from (`cube.provenance("forest_pct")` lists them per year).

DisSCube does not need DisSModel. To hand a cube to a DisSModel model, install
`disscube[dissmodel]` and use `cube.to_raster_backend(...)`, which returns a `RasterBackend`.

## Pipeline files (TOML)

A whole data preparation — grid, sources, derived variables — can be declared
in one TOML file, the counterpart of a TerraME `fill` script, and run from the
command line:

```bash
disscube validate examples/pipelines/quickstart.toml   # no downloads
disscube run      examples/pipelines/quickstart.toml --workspace outputs/quickstart
```

```toml
schema = 1
name = "Quickstart"

[grid]
name = "demo/300m"
crs = "EPSG:31983"
bbox = [570000.0, 9708000.0, 582000.0, 9720000.0]
resolution = 300

[[source]]
id = "landuse"
type = "file"
path = "../data/quickstart/landuse.tif"
crs = "EPSG:31983"

[[derive]]
target = "forest_pct"
source = "landuse"
operator = "percentage"
class_code = 3
```

See [`docs/guides/pipeline_files.md`](docs/guides/pipeline_files.md) and
[`examples/pipelines/`](examples/pipelines/).

## Examples

[`examples/`](examples/) has runnable, self-contained examples that run offline in seconds:

- **01–03, synthetic data** — a quickstart with raster operators, vector drivers, and time series handed off to DisSModel.
- **quickstart.toml** — declarative pipeline equivalent to example 01.

```bash
python examples/01_quickstart.py
disscube run examples/pipelines/quickstart.toml
```

Real data workflows, TerraME parity benchmarks, and large-scale case studies are maintained in
[**LambdaGeo/disscube-recipes**](https://github.com/LambdaGeo/disscube-recipes) (`cases/terrame_fill`, `cases/ilha_maranhao`, `cases/prodes_br163`, `cases/luccme_br`).

See [`examples/README.md`](examples/README.md) for the full list, and
[`docs/terrame_fill_correspondence.md`](docs/terrame_fill_correspondence.md)
for how DisSCube's operators relate to TerraME's *Fill*.

## Available operators

| Operator | Type | Resampling | Requires `class_code` |
|---|---|---|---|
| `mean` | zonal | average | no |
| `sum` | zonal | sum | no |
| `std` | zonal | nearest | no |
| `min` | zonal | min | no |
| `max` | zonal | max | no |
| `majority` | zonal | nearest¹ | no |
| `minority` | zonal | nearest¹ | no |
| `percentage` | zonal | nearest¹ | **yes** |
| `attribute` | zonal | nearest | no |
| `presence` | zonal | nearest | no |
| `distance` | proximity (exact, cell centre → nearest feature, source not clipped) | — | no |
| `min_distance` | proximity (raster approximation, features inside the grid) | nearest | no |
| `count` | proximity | nearest | no |
| `area` | polygons (exact share of the cell covered) | — | no |

`distance` takes `params = {crs = …}` to measure in another CRS (metres on a
geographic grid); the ¹ operators take `params = {subcells = n}` to cap the
fine pixels per cell. Any derivation takes `fill = "nearest"`.

> ¹ These use `needs_fine_alignment=True`: GridAligner resamples with `nearest` at high resolution; the actual reduction (per-window counting) is done by the operator.

## Pipeline

```
SpatialSource
    │
    ▼
Normalizer        — validates / loads a GeoDataFrame (vector) or opens the raster
    │
    ▼
GridAligner       — reprojects per variable with the operator's Resampling
    │
    ▼
Aggregator        — delegates to operator.compute() → one xr.DataArray per variable
    │
    ▼
VariableWriter    — writes Zarr + registers the DerivedVariable in the catalog
```

## Storage layout

```
data/derived/{grid_id}/{partition}/{spec_hash}/{variable_name}.zarr
```

- `partition` = `tile_id`, or `global` for untiled derivations.
- `spec_hash` = SHA-256 of the derivation (source + grid + variables + time window, plus the source's `checksum` when it has one — replacing a source file and registering its new checksum recomputes instead of returning a stale product).

## Project structure

```
disscube/
├── client.py         CubeClient — public entry point
├── models/           GridSpec, SpatialSource, SpatialDerivation, Variable, Derivation…
├── operators/        Operators as classes (self-registered via __init_subclass__)
│   ├── base.py       Operator ABC + OPERATOR_REGISTRY
│   ├── zonal.py      mean, sum, majority, percentage, attribute, presence…
│   └── proximity.py  distance, min_distance, count
├── pipeline/         Pipeline execution & planning (schema, runner) + internal stages
├── catalog/          CatalogStore (Protocol) + SQLite and JSON implementations
├── storage.py        AssetStore (fsspec — local and S3)
├── cli.py            `disscube validate` / `disscube run` / `disscube export`
├── sources/          Adapters that bring external data in as SpatialSources,
│   │                 each with a checksum and a provenance.json sidecar
│   ├── _raster.py    Window2D, windowed reads, composites, mosaics, register_raster
│   ├── bdc.py        Brazil Data Cube cubes via STAC (`bdc` extra)
│   ├── _categorical.py  legends (.qml/.json/.csv) and reclassification
│   ├── mapbiomas.py  MapBiomas annual land-cover maps (Collection 11, 10 m series)
│   ├── prodes.py     PRODES deforestation (download + cache, legend from the .qml)
│   └── classified.py any classified map with its legend, e.g. from SITS
└── utils.py          Checksums (sha256_file) and BDC tile importer (import_bdc_grids)
```

## Adding a new operator

Create a subclass of `Operator` in any module imported at startup:

```python
from rasterio.warp import Resampling
from disscube.operators.base import Operator

class WeightedMeanOperator(Operator):
    name = "weighted_mean"
    _resampling = Resampling.average

    def compute(self, data, var, grid):
        # data is an xr.DataArray (raster) or a GeoDataFrame (vector)
        ...
```

The operator is registered automatically and accepted by `Derivation` / `SpatialDerivation` with no other change.

## Known limitations

The limitations below are scope decisions for the current version, not bugs. They are documented so that users and reviewers understand what is implemented versus what is planned.

**In-memory, single-tile processing**
Each call to `derive()` loads a tile's full data into memory. There is no lazy (Dask) or distributed processing. For continental-scale grids (e.g. `BR/1km`), use the tile loop — each tile is processed and saved independently.

**Vector aggregation by rasterization (not area-weighted)**
Operators over vector sources (`majority`, `percentage`, `attribute`, `presence`, `minority`) convert geometries to raster before aggregating pixels. For the share of each cell covered by polygons, use `area`, which intersects the polygons with the cells exactly.

**Tile disambiguation in `load()`**
`CubeClient.load(name)` without `tile_id` raises `ValueError` when multiple tiles of the same variable exist on the same grid. Automatic mosaicking is not implemented. **Always pass `tile_id` in multi-tile workloads.**

**`SpatialRelation` does not act in the pipeline**
The `SpatialRelation` model is persisted in the catalog, but no pipeline stage uses relations during derivation — which is why they are **excluded from `spec_hash`**. Including them would make the cache key sensitive to metadata that does not affect the result, breaking the reproducibility guarantee. Integration with hierarchical grid strategies is reserved for a future version.

**`purity_threshold` reserved**
The `purity_threshold` field on `Derivation` is included in `spec_hash` but is not applied to the output — purity masking is not implemented. Setting `purity_threshold` changes the cache key without changing the result.

**STAC: reading only**
`disscube.sources.bdc` reads Brazil Data Cube cubes through their STAC catalog (search, windowed reads, per-tile composites, mosaics) and writes local GeoTIFFs that are registered as ordinary sources. Derived variables are not published back as STAC, and the `valid_from`/`valid_until` and `bbox` fields on `Derivation` only follow STAC naming conventions.

## Citation

If you use DisSCube in your research, dynamic modeling, or spatial data pipelines, please cite it using the metadata from [`CITATION.cff`](CITATION.cff) or the following BibTeX entry:

```bibtex
@software{costa_disscube_2026,
  author       = {Costa, S{\'e}rgio Souza},
  title        = {{DisSCube: Declarative spatial data cubes}},
  year         = {2026},
  version      = {0.4.0},
  url          = {https://github.com/DisSModel/disscube}
}
```

## License

DisSCube is part of the DisSModel ecosystem and is released under the MIT License. See [LICENSE](LICENSE) for details.
