Metadata-Version: 2.4
Name: dreamdb
Version: 0.0.15
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: License :: OSI Approved :: MIT License
License-File: LICENSE-MIT
Summary: Multimodal versioned data lake for ML training, on DreamDB
License-Expression: MIT
Requires-Python: >=3.8
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://dreamdb.dreamlake.ai
Project-URL: Issues, https://github.com/dreamlake-ai/dreamdb-core/issues
Project-URL: Repository, https://github.com/dreamlake-ai/dreamdb-core

# DreamDB Python SDK

Python bindings for the [DreamDB](https://github.com/dreamlake-ai/dreamdb-core) multimodal
versioned data lake — image + audio + text + embeddings + scalar
metadata on content-addressed object storage, with vector and
metadata filters for ML training pipelines.

## Quick start

The SDK exposes read-only `SqlSession.open(ref_name, backend)` with
`query` and `explain`. See [SQL adapters](../dreamdb-sql/ADAPTERS.md) for typed
parameters, snapshot semantics and limits.

```python
from pathlib import Path
from tempfile import TemporaryDirectory
import dreamdb as db

with TemporaryDirectory(prefix="dreamdb-quickstart-") as directory:
    backend = Path(directory).as_uri()
    schema = db.Schema().add_scalar_string("label")
    ds = db.Dataset.create("example", schema, backend=backend)
    ds.append_many([{"_anchor": 1, "label": "cat"}])
    reopened = db.Dataset.open("example", backend=backend)
    assert reopened.count() == 1
```

For trained vector indexes, use the [versioned index guide](https://dreamdb.dreamlake.ai/python-sdk-indexes).
The [feature examples](../docs/main-features.md) cover Python 0.0.11, including
entity keys, progressive geometry and structured arrays. Older packages do not
necessarily expose these methods.

## Build from source

### Segmented videos with independent initialization (Python 0.0.15)

`Schema.add_video_item("video_raw", codec="h265")` declares optional VideoItems.
Unlike a flat CMAF Track, each item owns its decoder initialization and relative
fragment index. The application prepares the media (for example by stream-copy
remuxing at source keyframes); DreamDB does not transcode or create previews.

```python
schema = db.Schema().add_video_item("video_raw", codec="h265")
ds = db.Dataset.create("new-originals", schema, backend=backend)
# fragments: [(fragment_bytes, relative_start_ns, relative_end_ns), ...]
published = ds.publish_prepared_video_item(
    "video_raw", b"opaque-clip-key", absolute_start_ns, duration_ns,
    init_bytes, fragments,
)
reopened = db.Dataset.open("new-originals", backend=backend)
window = reopened.read_video_item_range("video_raw", b"opaque-clip-key", 0, duration_ns)
```

This API materializes one item's inputs in memory, copies them into native
ownership and publishes through the existing core validation/CAS boundary.
It is not streaming, zero-copy, or a browser compatibility guarantee. Reusing
an item key replaces that item; concurrent publication can raise a CAS conflict
and is not implicitly retried. Returned addresses identify item, Track and
Manifest. These write/schema methods are included in Python 0.0.15; 0.0.14
does not expose them.

### Building the bindings

```bash
pip install maturin
cd dreamdb-dataset-python
maturin develop --release
```

This produces the `dreamdb` package installed into the
current virtualenv. Importing it gives the `Schema` and `Dataset`
classes shown above.
# Publisher identity

New Manifest publications record the Python distribution version, shared core
build identity and operation in the existing `writer` tag. Optional
`dataset.set_application(name, revision)` identifies your application on later
publications; reopening a Dataset resets that optional pair. This does not label
retained payloads or prove execution. See [publisher provenance](../docs/publisher-provenance.md)
for limits, build configuration and historical/operator behavior.
# Explicit ingest checkpoints

For long-running **base scalar + VideoItem** ingest, a caller may explicitly
start a new provenance epoch while reusing the current data objects:

```python
used, limit = dataset.ingest_lineage_usage()
plan = dataset.plan_ingest_checkpoint(backend, "ingest-epoch-001")
# Durably persist these exact bytes BEFORE apply (file flush + fsync in your
# application's checkpoint journal); never rebuild a plan after an uncertain PUT.
save_durably(plan)
tip = dataset.apply_ingest_checkpoint(plan, backend)
```

`backend` here is the same non-secret endpoint/bucket identity on both calls.
Exclude concurrent GC. The operation pins the old tip with a create-only tag,
publishes a frozen new root and conditionally advances the working Ref. After
an uncertain result, reopen that Ref and apply the **same saved plan**. A third
tip conflicts; there is no automatic merge. A successful application preserves
current data, not its provenance chain: old history remains under the tag, and
pre-checkpoint branches cannot automatically merge into the new epoch.

Embeddings and layers are rejected by the base-only planner. For **cold IVF-cosine
embeddings, including layers over VideoItem**, use
`plan_indexed_ingest_checkpoint(backend, archive_tag)` and persist/apply the plan
the same way. It retains exact original source bindings as structural ancestry
claims, preserves SI/compressor/sidecars and current roles, and emits **lineage-v5**.
All readers/writers must support that format before opting in; older packages
reject it. Merge/backfill/entity combinations and unknown dependency-bearing
metadata remain refused. This does not automatically publish new SDK packages.

Indexed admission reads bounded Track roots (4 MiB each, 32 MiB aggregate), not
all bucket payloads; oversized imported roots need explicit conversion. It does
not certify imported payload completeness. Existing full-directory compaction
is separate, not implicit checkpoint work. See [spec 0030](../spec/0030-continuous-index-maintenance.md).

The usage result is a measurement, not a guarantee that
the next append fits. No object deletion, GC, tag expiry, or larger protocol
limit is implied. See [spec 0029](../spec/0029-ingest-checkpoints.md).

