Metadata-Version: 2.4
Name: hivemind-sdk
Version: 0.2.2
Summary: Client SDK for HiveMind: persistent shared memory and a working-context compiler for agent systems.
Author-email: "Militant.AI" <support@militant.ai>
License: Proprietary
Project-URL: Homepage, https://hivemind.militant.ai
Project-URL: Documentation, https://hivemind.militant.ai/docs
Keywords: ai,agents,memory,llm,context,rag
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: httpx>=0.24

# HiveMind Python SDK

Persistent memory and a working-context compiler for your agent, in three
lines inside your own loop. Your model, your key — HiveMind never calls an
LLM.

```bash
pip install hivemind-sdk
```

Python ≥ 3.10. The import name is `hivemind`.
Docs: https://hivemind.militant.ai/docs

## Async

Agent systems run on the event loop; so does this SDK. `AsyncHiveMind`
mirrors the sync facade — same three-liner, nothing blocks, pooled
connections underneath:

```python
from hivemind import AsyncHiveMind

async with AsyncHiveMind(base_url="...", api_key="...") as mind:
    session = mind.session(budget_total=8192)
    result = await session.turn(user_input)
    reply = await call_your_llm(result.messages)
    await session.record(reply)
```

`AsyncHivemindClient` underneath mirrors the engine's native
`HivemindOperations` surface name-for-name (`compile_working_context`,
`store_conversation_exchange`, `recall_memory_with_metadata`, …) — code
written against the ops layer speaks to the hosted service with the same
vocabulary. The sync client remains pure standard library.

```python
from hivemind import HiveMind

mind = HiveMind(base_url="...", api_key="...", tenant_id="...")
session = mind.session(budget_total=8192)

result = session.turn(user_input)       # store -> recall -> compile
reply = call_your_llm(result.messages)  # your model, your key
session.record(reply)                   # completes the exchange
```

## What one `turn()` does

One HTTP request (`POST /turn`) — the service composes, in order:

1. Stores the user message (receipted).
2. Semantically recalls relevant memories and past conversation.
3. Compiles local history + recalled records + active holds + your
   operator briefing into a token-budgeted bundle.

With `record()`, the whole exchange is two calls — your model runs
between them.

`result.messages` is the **entire** prompt payload — send it as-is, splice
nothing in. `result.bundle["decisions"]` explains every admission under
the budget; `result.receipt` is the audit record. Empty recall on a young
tenant is normal, not an error.

`session.record(reply)` stores your model's reply as the other half of the
exchange, so the next turn — and every future session — remembers it.

## How conversation memory recalls

Conversation is remembered as **call/response exchanges** — a user
question, an agent's instruction, whatever the initiating text was, plus
the reply it produced. Recall matches your query against **both sides**
of every past exchange, and returns **whole exchanges**: one result slot
per exchange, rendered call-then-response, never an answer without the
message that produced it (and vice versa). Facts that appear only in a
reply are just as findable as the calls that prompted them.

The recall pool is sized automatically from your session's token budget —
a bigger `budget_total` recalls more candidates, and the compiler's
budget admission decides what actually enters the bundle (with every
decision receipted). Pass `recall_top_k` to a session only if you want to
force a fixed pool.

Still worth designing around: `record()` files the reply and completes
the exchange — treat it as part of the loop. And durable facts that
should stand alone — decisions, outcomes, lessons — belong in
`mind.remember(...)`, where you control their metadata and lifecycle.

Every turn also reports where its time went: `result.timings` carries the
server-side phase breakdown in seconds (`embed_s`, `store_and_recall_s`,
`completion_s`, `compile_s`, `total_s`) — a slow turn names its own
bottleneck.

## Beyond the loop

- `mind.remember(content, metadata)` — deliberately store a durable
  lesson, decision, fact, or outcome.
- `mind.recall(query)` / `mind.recall_filtered(query, metadata)` —
  explicit recall over **deliberate memories only** (`[]` when nothing
  matches). Conversation history is a separate record class, recalled
  automatically inside the loop — or explicitly via
  `client.recall_conversation`.
- `session.hold_set(key, content)` / `hold_clear(key)` — pin operational
  state ("stop-order", "API is down") into every compile until cleared.
- `mind.receipts(session_id=...)` — the audit trail: what ran, what it
  consumed, what it produced, with lineage.
- `mind.delete_by_metadata(metadata)` — destructive, audited deletion.
- `mind.client` — the raw HTTP client for anything not wrapped.

## Configuration

Constructor arguments override environment:

| Env var | Meaning |
|---|---|
| `HIVEMIND_BASE_URL` | Service root (hosted or local — same API) |
| `HIVEMIND_API_KEY` | Sent as `Authorization: Bearer <key>` |
| `HIVEMIND_TENANT_ID` | Your tenant (`X-Tenant-ID`) |
| `HIVEMIND_TIMEOUT` | Request timeout, seconds (default 30) |
| `HIVEMIND_BUDGET_TOTAL` | Default compile token budget (default 4096) |

## Development note (this repo)

The import name `hivemind` collides with the service package at the repo
root, so run SDK tests as their own invocation:

```bash
python -m pytest sdk/python/tests
```
