Metadata-Version: 2.4
Name: cairndb
Version: 0.4.1
Summary: Serverless database engine on blob storage: commit logs, coordination primitives, multi-key transactions, and SQLite projections — no servers anywhere
Author: CairnDB Contributors
License-Expression: MIT
Project-URL: Documentation, https://quadratic-labs.github.io/cairndb/
Project-URL: Repository, https://github.com/Quadratic-Labs/cairndb
Project-URL: Changelog, https://github.com/Quadratic-Labs/cairndb/blob/main/CHANGELOG.md
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.14
Requires-Python: >=3.14
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: msgpack>=1.0.0
Requires-Dist: sqlalchemy[asyncio]>=2.0.25
Requires-Dist: aiosqlite>=0.19.0
Requires-Dist: structlog>=24.1.0
Provides-Extra: s3
Requires-Dist: boto3>=1.35.0; extra == "s3"
Provides-Extra: gcs
Requires-Dist: google-cloud-storage>=2.14.0; extra == "gcs"
Provides-Extra: azure
Requires-Dist: azure-storage-blob>=12.19.0; extra == "azure"
Requires-Dist: azure-identity>=1.16.0; extra == "azure"
Provides-Extra: cli
Requires-Dist: typer>=0.9.0; extra == "cli"
Provides-Extra: docs
Requires-Dist: sphinx>=8.0; extra == "docs"
Requires-Dist: myst-parser>=4.0; extra == "docs"
Requires-Dist: furo>=2024.8.6; extra == "docs"
Requires-Dist: sphinx-copybutton>=0.5.2; extra == "docs"
Requires-Dist: sphinxcontrib-mermaid>=1.0; extra == "docs"
Requires-Dist: sphinx-autobuild>=2024.10.3; extra == "docs"
Requires-Dist: typer>=0.9.0; extra == "docs"
Provides-Extra: dev
Requires-Dist: pytest>=8.0.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23.0; extra == "dev"
Requires-Dist: pytest-cov>=4.1.0; extra == "dev"
Requires-Dist: hypothesis>=6.98.0; extra == "dev"
Requires-Dist: mutmut>=3.7.0; extra == "dev"
Requires-Dist: black>=24.1.0; extra == "dev"
Requires-Dist: ruff>=0.2.0; extra == "dev"
Requires-Dist: mypy>=1.8.0; extra == "dev"
Requires-Dist: boto3>=1.35.0; extra == "dev"
Requires-Dist: google-cloud-storage>=2.14.0; extra == "dev"
Requires-Dist: azure-storage-blob>=12.19.0; extra == "dev"
Requires-Dist: azure-identity>=1.16.0; extra == "dev"
Requires-Dist: typer>=0.9.0; extra == "dev"
Dynamic: license-file

# CairnDB

**A serverless database engine on blob storage — the bucket is the only server.**

A cairn is a stack of stones raised one by one, by many independent hands,
with no custodian — and it stands for centuries. CairnDB works the same
way: all state lives in an object-storage bucket, every writer is just a
library call, and the bucket itself arbitrates concurrency through
conditional writes. No database server, no coordinator, no idle compute.

On that substrate CairnDB exposes the layers a database engine is made of:

| Layer | You get | Engine analogy |
|---|---|---|
| Objects | etag-guarded key-value documents (CAS, put-if-absent) | atomic page writes |
| Coordination | `claim` (unique constraint), `lease` (fenced ownership), `doc` (retrying read-modify-write) | locks & constraints |
| Logs | named, append-only commit logs — dense, totally ordered, durable on ack | WAL |
| Transactions | optimistic multi-key atomicity, coordinated by a system log | transaction manager |
| Projections | deterministic replay into local read-only SQLite, snapshots, time travel | indexes & materialized views |

```python
from cairndb import CairnDB

db = CairnDB.configure({"storage": {"type": "s3", "bucket": "myapp"}})

await db.objects.put("config/app.json", data)                  # conditional KV
result = await db.claim("dispatch/etl:2026-08-10", {"run": 1})  # exactly-one winner
lease  = await db.lease("state/run-1", ttl=120)                 # fenced ownership
seq    = await db.log("orders").append(event)                   # durable, ordered
async with db.transact() as tx:                                 # multi-key atomicity
    tx.put("accounts/alice", alice_bytes)
    tx.put("accounts/bob", bob_bytes)
proj   = db.projection("orders_view", log="orders")             # SQL over the log
```

The documentation is published at **<https://quadratic-labs.github.io/cairndb/>**
(sources in [`docs/`](https://github.com/Quadratic-Labs/cairndb/tree/main/docs)). Start with
the [quickstart](https://quadratic-labs.github.io/cairndb/getting-started/quickstart.html),
and see the [concepts](https://quadratic-labs.github.io/cairndb/concepts/index.html) for the
semantics of each layer.

## Use Cases

- Small to medium datasets (up to ~10 GB per projection) with complex read
  queries and low/medium write volume
- Strong auditability and determinism requirements: event sourcing,
  time-travel reads, rebuild-anywhere recovery
- Coordination state for serverless/scale-to-zero systems: workflow
  ownership, exactly-once dispatch, checkpoints
- "I want a database but refuse to run or rent a database server"

## How It Works

```
writers ──PUT log/N+1 (if-absent)──▶  ┌────────────────────────┐
                                      │  Blob storage bucket   │
readers ◀──GET log/N+1 (poll)──────── │  log/000000000042.msgpack
                                      │  logs/{name}/...        │
                                      │  snapshots/v1/...sqlite │
snapshot job (cron) ◀──replay──────▶  └────────────────────────┘
```

1. Every write batch becomes an immutable, numbered commit object. The
   bucket accepts exactly one writer per number (`If-None-Match`), so each
   log is dense, gap-free, and totally ordered — no clocks, no sequencer
   service. Coordination documents use the same conditional-write
   machinery with etags (compare-and-swap).
2. Clients replay commits through registered event handlers into a local
   SQLite file, swapped atomically so readers never see partial state.
3. A scheduled job replays the log into snapshot files that bound client
   startup time; garbage collection prunes what snapshots cover. Nothing
   outside the bucket needs to survive.

Multiple concurrent writers are safe by construction: a writer that loses
the race for commit N+1 fetches the winner's commits, optionally
revalidates its events against them, and retries at N+2. A fenced lease
holder's writes are rejected by the storage itself.

## Installation

```bash
# Requires Python 3.14+
pip install cairndb            # filesystem backend only
pip install cairndb[s3]        # + Amazon S3 / S3-compatible
pip install cairndb[gcs]       # + Google Cloud Storage
pip install cairndb[azure]     # + Azure Blob Storage
pip install cairndb[cli]       # + `cairndb` CLI (snapshot/gc/rebuild jobs)
```

## Quick Start

```python
import aiosqlite
from cairndb import CairnDB, Event, EventType, SchemaVersion, Timestamp

async def init_users(db_path: str) -> None:
    async with aiosqlite.connect(db_path) as conn:
        await conn.execute("CREATE TABLE IF NOT EXISTS users (id INTEGER PRIMARY KEY, name TEXT)")
        await conn.commit()

db = CairnDB.configure({"storage": {"type": "filesystem", "path": "./data"}})

# Write: durable and totally ordered as soon as append() returns
seq = await db.log("users").append(
    Event(
        event_type=EventType("user.created"),
        timestamp=Timestamp.now(),
        payload={"id": 1, "name": "Alice"},
        schema_version=SchemaVersion("1"),
    )
)

# Project: replay events into local SQLite, declaratively
proj = db.projection("users_view", log="users", init_schema=init_users)

@proj.on("user.created")
async def handle_user_created(conn, entry):
    await conn.execute(
        "INSERT INTO users (id, name) VALUES (?, ?)",
        (entry.payload["id"], entry.payload["name"]),
    )

await proj.wait_for(seq)                     # read-your-writes
with proj.connect() as conn:                 # read-only SQLite
    rows = conn.execute("SELECT name FROM users").fetchall()

await db.close()
```

Coordination needs no handlers or projections — it is storage-level:

```python
result = await db.claim(f"dispatch/{key}", {"run_id": run_id})
if result.won:
    lease = await db.lease(f"state/{run_id}", ttl=120, holder=worker_id)
    ...work, calling await lease.renew() as a heartbeat...
    await lease.release(state={"status": "done"})   # LeaseLost if we were fenced
```

### Snapshots & retention (scheduled jobs)

```bash
export CAIRNDB_STORAGE_TYPE=s3 CAIRNDB_S3_BUCKET=my-bucket

cairndb snapshot --handlers myapp.projections:registry \
                  --init-schema myapp.projections:init_schema
cairndb snapshot --log orders --handlers myapp.projections:orders_registry
cairndb gc --keep-snapshots 3            # add --prune-log to drop covered history
```

See the [quickstart](https://quadratic-labs.github.io/cairndb/getting-started/quickstart.html) for the full walkthrough,
including the lower-level `Committer`/`HandlerRegistry` API the facade is
built on.

## Guarantees & Trade-offs

| Property | Guarantee |
|---|---|
| Durability | An acked write is in the bucket; there is no ack-before-durable window |
| Ordering | Total order per log, arbitrated by the bucket (no clocks) |
| Concurrency | Any number of writers; optimistic races with retry; epoch-fenced leases |
| Transactions | Optimistic multi-key atomicity; commit point is a log record; conflicting transactions abort |
| Read consistency | Eventual (poll interval); per-session read-your-writes via `wait_for` |
| Throughput ceiling | ~1 commit per storage round-trip per log (batching multiplies events/commit; shard across named logs) |
| Recovery | Everything outside the bucket is disposable and rebuilt by replay |

## License

[MIT](https://github.com/Quadratic-Labs/cairndb/blob/main/LICENSE) © 2026 Thomas Zamojski
