Metadata-Version: 2.4
Name: gestaltdb
Version: 0.3.0
Summary: A pure Python GraphDB for attributed graphs.
Requires-Python: <3.14,>=3.9
Description-Content-Type: text/markdown
Requires-Dist: google>=3.0.0
Requires-Dist: msgpack>=1.1.2
Requires-Dist: plyvel>=1.5.1
Requires-Dist: protobuf>=6.33.6
Provides-Extra: lmdb
Requires-Dist: lmdb; extra == "lmdb"
Provides-Extra: leveldb
Requires-Dist: plyvel; extra == "leveldb"
Provides-Extra: msgpack
Requires-Dist: msgpack; extra == "msgpack"
Provides-Extra: protobuf
Requires-Dist: protobuf; extra == "protobuf"
Provides-Extra: bloom
Requires-Dist: pybloom-live; extra == "bloom"
Provides-Extra: coverage
Requires-Dist: coverage; extra == "coverage"
Requires-Dist: pytest; extra == "coverage"
Provides-Extra: docs
Requires-Dist: furo; extra == "docs"
Requires-Dist: myst-parser; extra == "docs"
Requires-Dist: sphinx; extra == "docs"
Provides-Extra: rocksdb
Requires-Dist: pyrex-rocksdb>=0.3.0a0; extra == "rocksdb"
Provides-Extra: arrow
Requires-Dist: pyarrow; extra == "arrow"
Provides-Extra: polars
Requires-Dist: polars; extra == "polars"
Requires-Dist: pyarrow; extra == "polars"
Provides-Extra: fast-ingest
Requires-Dist: pyarrow; extra == "fast-ingest"
Requires-Dist: polars; extra == "fast-ingest"
Requires-Dist: pyrex-rocksdb>=0.3.0a0; extra == "fast-ingest"
Provides-Extra: all
Requires-Dist: lmdb; extra == "all"
Requires-Dist: msgpack; extra == "all"
Requires-Dist: plyvel; extra == "all"
Requires-Dist: protobuf; extra == "all"
Requires-Dist: pybloom-live; extra == "all"
Requires-Dist: pyrex-rocksdb>=0.3.0a0; extra == "all"
Requires-Dist: pyarrow; extra == "all"
Requires-Dist: polars; extra == "all"
Provides-Extra: dev
Requires-Dist: coverage; extra == "dev"
Requires-Dist: furo; extra == "dev"
Requires-Dist: lmdb; extra == "dev"
Requires-Dist: msgpack; extra == "dev"
Requires-Dist: myst-parser; extra == "dev"
Requires-Dist: plyvel; extra == "dev"
Requires-Dist: protobuf; extra == "dev"
Requires-Dist: pybloom-live; extra == "dev"
Requires-Dist: pyrex-rocksdb>=0.3.0a0; extra == "dev"
Requires-Dist: pyarrow; extra == "dev"
Requires-Dist: polars; extra == "dev"
Requires-Dist: pytest; extra == "dev"
Requires-Dist: sphinx; extra == "dev"

# GestaltDB

![Coverage](https://raw.githubusercontent.com/mylonasc/gestaltdb/refs/heads/main/assets/coverage_badge.svg)
[![Documentation](https://img.shields.io/badge/docs-GitHub%20Pages-blue.svg)](https://mylonasc.github.io/gestaltdb/)

GestaltDB is a pure Python graph database toolkit for attributed graphs. It stores nodes, edges, labels, typed adjacency records, and property indexes on embedded key-value backends.

Documentation: https://mylonasc.github.io/gestaltdb/

## Install From PyPI

With pip:

```sh
python -m pip install gestaltdb
```

With uv:

```sh
uv add gestaltdb
```

Install columnar ingestion dependencies:

```sh
python -m pip install "gestaltdb[arrow,polars]"
```

Install all optional backends and serializers:

```sh
python -m pip install "gestaltdb[all]"
```

Optional extras include `lmdb`, `leveldb`, `rocksdb`, `arrow`, `polars`, `fast-ingest`, `msgpack`, `protobuf`, `bloom`, `docs`, `dev`, and `all`.

## Basic Example

```python
from tempfile import TemporaryDirectory

from gestaltdb.graphdb import Edge, GraphDB, Node
from gestaltdb.kvstores import LevelDBStore
from gestaltdb.serializers import PickleSerializer

with TemporaryDirectory() as tmpdir:
    graph = GraphDB(LevelDBStore(path=f"{tmpdir}/graph"), PickleSerializer())

    graph.put_node(Node(node_id="alice", labels=["Person"], properties={"name": "Alice"}))
    graph.put_node(Node(node_id="bob", labels=["Person"], properties={"name": "Bob"}))
    graph.put_edge(Edge(
        edge_id="alice-knows-bob",
        source="alice",
        target="bob",
        properties={"type": "knows", "since": 2024},
    ))

    result = graph.query('MATCH (a:Person {name: "Alice"}) MATCH (a)-[:knows]->(b) RETURN a.id, b.name')
    print(result.records)

    graph.close()
```

## Arrow Ingestion Example

This example ingests entity columns from PyArrow arrays. `JSONSerializer` lets GestaltDB build node and edge payloads from structured columns.

```python
from tempfile import TemporaryDirectory

import pyarrow as pa

from gestaltdb.graphdb import GraphDB
from gestaltdb.kvstores import LevelDBStore
from gestaltdb.serializers import JSONSerializer

with TemporaryDirectory() as tmpdir:
    graph = GraphDB(LevelDBStore(path=f"{tmpdir}/graph"), JSONSerializer())

    graph.ingest_nodes_arrow_entities(
        pa.array(["alice", "bob", "carol"]),
        labels=pa.array([["Person"], ["Person"], ["Person"]]),
        properties={
            "name": pa.array(["Alice", "Bob", "Carol"]),
            "age": pa.array([34, 36, 29]),
        },
    )

    graph.ingest_edges_arrow_entities(
        pa.array(["alice-knows-bob", "bob-knows-carol"]),
        pa.array(["alice", "bob"]),
        pa.array(["bob", "carol"]),
        pa.array(["knows", "knows"]),
        properties={"since": pa.array([2024, 2025])},
    )

    result = graph.query('MATCH (a:Person {name: "Alice"}) MATCH (a)-[:knows]->(b) RETURN a.id, b.name')
    print(result.records)

    graph.close()
```

## Polars Ingestion Example

This example ingests the same graph from Polars DataFrames. Property columns are converted into node and edge payloads during ingestion.

```python
from tempfile import TemporaryDirectory

import polars as pl

from gestaltdb.graphdb import GraphDB
from gestaltdb.kvstores import LevelDBStore
from gestaltdb.serializers import JSONSerializer

nodes = pl.DataFrame({
    "node_id": ["alice", "bob", "carol"],
    "labels": [["Person"], ["Person"], ["Person"]],
    "name": ["Alice", "Bob", "Carol"],
    "age": [34, 36, 29],
})

edges = pl.DataFrame({
    "edge_id": ["alice-knows-bob", "bob-knows-carol"],
    "source": ["alice", "bob"],
    "target": ["bob", "carol"],
    "edge_type": ["knows", "knows"],
    "since": [2024, 2025],
})

with TemporaryDirectory() as tmpdir:
    graph = GraphDB(LevelDBStore(path=f"{tmpdir}/graph"), JSONSerializer())

    graph.ingest_nodes_polars_entities(nodes)
    graph.ingest_edges_polars_entities(edges)

    result = graph.query('MATCH (a:Person) MATCH (a)-[:knows]->(b) RETURN a.name, b.name ORDER BY a.name')
    print(result.records)

    graph.close()
```

## Install From A Checkout

From a local checkout:

```sh
uv sync
```

Install into another project:

```sh
uv add /path/to/gestaltdb
```

With pip:

```sh
python -m pip install /path/to/gestaltdb
```

## Features

- Attributed `Node` and `Edge` objects with stable IDs.
- Native node labels and typed edge traversal through `edge.properties["type"]`.
- LMDB, LevelDB, and RocksDB/PyRex storage backends.
- Pickle, JSON, MessagePack, and Protobuf serializers.
- Label, relationship type, property, composite, and range indexes.
- Read-only Cypher subset for indexed scans, typed traversal, filtering, ordering, limits, and chained `MATCH` clauses.
- Bulk and columnar ingestion helpers for Arrow and Polars.
- Typed path and subgraph sampling.

See the full documentation for backend selection, indexing, Cypher syntax, ingestion, sampling, and benchmarks.

<details>
<summary>Name origin</summary>

The name GestaltDB is inspired by Gestalt psychology and the idea that the whole is something more than its parts.

</details>
