Metadata-Version: 2.4
Name: dreamdb
Version: 0.0.12
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: License :: OSI Approved :: MIT License
License-File: LICENSE-MIT
Summary: Multimodal versioned data lake for ML training, on DreamDB
License-Expression: MIT
Requires-Python: >=3.8
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://dreamdb.dreamlake.ai
Project-URL: Issues, https://github.com/dreamlake-ai/dreamdb-core/issues
Project-URL: Repository, https://github.com/dreamlake-ai/dreamdb-core

# DreamDB Python SDK

Python bindings for the [DreamDB](https://github.com/dreamlake-ai/dreamdb-core) multimodal
versioned data lake — image + audio + text + embeddings + scalar
metadata on content-addressed object storage, with vector and
metadata filters for ML training pipelines.

## Quick start

```python
from pathlib import Path
from tempfile import TemporaryDirectory
import dreamdb as db

with TemporaryDirectory(prefix="dreamdb-quickstart-") as directory:
    backend = Path(directory).as_uri()
    schema = db.Schema().add_scalar_string("label")
    ds = db.Dataset.create("example", schema, backend=backend)
    ds.append_many([{"_anchor": 1, "label": "cat"}])
    reopened = db.Dataset.open("example", backend=backend)
    assert reopened.count() == 1
```

For trained vector indexes, use the [versioned index guide](https://dreamdb.dreamlake.ai/python-sdk-indexes).
The [feature examples](../docs/main-features.md) cover Python 0.0.11, including
entity keys, progressive geometry and structured arrays. Older packages do not
necessarily expose these methods.

## Build from source

```bash
pip install maturin
cd dreamdb-dataset-python
maturin develop --release
```

This produces the `dreamdb` package installed into the
current virtualenv. Importing it gives the `Schema` and `Dataset`
classes shown above.

