Metadata-Version: 2.4
Name: zarr-beluga
Version: 0.1.0
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: 3
Classifier: License :: OSI Approved :: MIT License
Classifier: Topic :: Scientific/Engineering :: Atmospheric Science
Requires-Dist: numpy>=1.23
Requires-Dist: torch>=2.1 ; extra == 'bench'
Requires-Dist: xarray>=2025.1 ; extra == 'bench'
Requires-Dist: zarr>=3.1 ; extra == 'bench'
Requires-Dist: obstore>=0.6 ; extra == 'bench'
Requires-Dist: fsspec ; extra == 'bench'
Requires-Dist: adlfs ; extra == 'bench'
Requires-Dist: dask ; extra == 'bench'
Requires-Dist: psutil ; extra == 'bench'
Requires-Dist: tomli ; python_full_version < '3.11' and extra == 'bench'
Requires-Dist: pytest ; extra == 'test'
Requires-Dist: torch>=2.1 ; extra == 'test'
Requires-Dist: xarray>=2025.1 ; extra == 'test'
Requires-Dist: zarr>=3.1 ; extra == 'test'
Requires-Dist: obstore>=0.6 ; extra == 'test'
Requires-Dist: dask ; extra == 'test'
Requires-Dist: torch>=2.1 ; extra == 'torch'
Provides-Extra: bench
Provides-Extra: test
Provides-Extra: torch
License-File: LICENSE
Summary: Beluga: a fast reader for random-over-time, dense-over-space weather data (zarr v2/v3, Beluga v0) from blob storage, for ML training
Keywords: zarr,weather,climate,era5,dataloader,pytorch
License: MIT
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM

# zarr-beluga

Beluga is a fast reader for one access pattern: random reads over time, dense
over space and variables, of gridded weather and climate data from blob
storage into GPU training loops.

It reads zarr v2 groups in the WeatherBench2 layout (one array per variable),
single-array zarr v3 stores (sharded or not) and Beluga's own format, from
local disk, Azure Blob, S3 or GCS. The reader is written in Rust: it keeps
many range requests in flight across sample boundaries, decodes in parallel
straight into the output buffer, and hands each sample to Python without a
copy (DLPack).

## Install

```bash
pip install zarr-beluga            # the import name is `beluga`
pip install "zarr-beluga[torch]"   # with the PyTorch dataset
```

Wheels are built for Linux x86_64 and macOS arm64, Python 3.10 and later.

## Use

```python
import beluga
import torch

reader = beluga.Reader(
    "az://my-container",
    root="era5.zarr",
    store="objstore",
    store_options={
        "azure_storage_account_name": "myaccount",
        "azure_storage_sas_key_file": "/path/to/era5.sas",
    },
    source="zarr-group",
    scheduler="zarrs-group",
)
print(reader.meta["channels"], reader.meta["n_time"])

# Windows of 4 consecutive timesteps starting at these time indices.
for sample in reader.stream([10, 3, 7], frames=4):
    x = torch.from_dlpack(sample)  # (4, channels, height, width) float32
```

A local directory works the same way (`beluga.Reader("/data", root="era5.zarr", ...)`).
Credentials can also come from the usual `AZURE_*`, `AWS_*` and `GOOGLE_*`
environment variables.

- `beluga.torch` has `BelugaDataset`, an `IterableDataset` that streams a
  shuffled epoch from one reader, and `DeviceLoader`, which moves samples to
  the GPU one step ahead.
- `beluga.otter.make_dataloader` is a drop-in for the data loaders of
  [otter](https://github.com/jonas-scholz123/otter)'s training configs.

