Metadata-Version: 2.5
Name: overmind
Version: 0.1.84
Summary: Overmind — autonomous agent optimisation through structured experimentation
Project-URL: Homepage, https://github.com/overmind-core/overmind
Project-URL: Repository, https://github.com/overmind-core/overmind
Author: Overmind Ltd
License: MIT
License-File: LICENSE
Keywords: ai,api,client,enforcement,overmind,policy
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Requires-Python: <4,>=3.10
Requires-Dist: posthog>=7.53.0
Requires-Dist: requests>=2.31.0
Requires-Dist: rich>=14.0.0
Requires-Dist: tomli-w>=1.0.0
Requires-Dist: typer>=0.25.0
Provides-Extra: tracing
Requires-Dist: openinference-instrumentation-langchain>=0.1.72; extra == 'tracing'
Requires-Dist: opentelemetry-api>=1.35.0; extra == 'tracing'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.35.0; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-agno>=0.53.0; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-anthropic>=0.53.0; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-fastapi>=0.60b1; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-google-generativeai>=0.53.0; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-httpx>=0.60b1; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-logging>=0.60b1; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-openai>=0.53.1; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation-requests>=0.60b1; extra == 'tracing'
Requires-Dist: opentelemetry-instrumentation>=0.60b1; extra == 'tracing'
Requires-Dist: opentelemetry-overmind>=0.53.0; extra == 'tracing'
Requires-Dist: opentelemetry-sdk>=1.35.0; extra == 'tracing'
Requires-Dist: opentelemetry-semantic-conventions-ai>=0.4.13; extra == 'tracing'
Requires-Dist: opentelemetry-semantic-conventions>=0.60b1; extra == 'tracing'
Description-Content-Type: text/markdown

<img alt="Overmind" src="https://github.com/user-attachments/assets/4a5caceb-49e8-4b8e-a6aa-511222a94381" />

# Overmind

Overmind is two things in one package:

- **Tracing SDK** — drop-in observability for LLM agents. Decorate your code, get structured traces of every LLM call and tool invocation.
- **Inference client and CLI** — `overmind.Client` calls models you deployed on Overmind through the OpenAI-compatible API; the `overmind` command scans, syncs, and moves datasets and checkpoints.

**Documentation:** [Overmind guide](https://docs.overmindlab.ai/core/observability)

**Console:** [console.overmindlab.ai](https://console.overmindlab.ai/)

## Install

```bash
pip install overmind              # CLI (`overmind init`, sync, dataset/model files, optimiser)
pip install "overmind[tracing]"   # OpenTelemetry tracing (optional extra)
```

The default install is the command-line tool. Tracing is a separate extra so an
app that already pinned OpenTelemetry does not clash with ours.

On Windows, install into a virtual environment; see
[Windows](https://docs.overmindlab.ai/platform/cli#windows).

```bash
uv tool install overmind
# or
pipx install overmind
```

## Quick start (local setup)

The optional plugin connects through OAuth and can access every project your
Overmind account is authorized to access. Start with `list_projects`, then pass
the chosen `project_id` on each project tool and resource URI. Account API keys
support the same access model. The local repository setup below intentionally
installs a narrower project API key for that repository.

```bash
export OVERMIND_API_KEY=<your-api-key>
export OVERMIND_API_URL=https://api.overmindlab.ai   # or your console API host

pip install overmind
overmind init --ide cursor   # or claude | opencode | codex
overmind sync
```

The pasted account key is a bootstrap credential for the current shell only.
`init` installs the skill and prepares the selected IDE; `sync` creates the
project, stores its project-scoped key in `.overmind/credentials.toml`, and
updates the local MCP config. Reload the IDE once after the first sync.

Then in your coding harness:

```text
/overmind setup              # scan → capabilities → evals → overmind.toml → sync
```

The installed skill opens with the detected repository, package manager, and
verified connection state, then previews the full onboarding roadmap: installation,
connection, and eight scan stages. Each conversation update names the current
stage, its number out of ten, how many stages are complete, and how many follow
it. Discovery explains how the identified capabilities fit together; subsequent
updates report repository-specific findings. Longer stages include per-capability counts. It reports the coding-agent
model when available, the Overmind destination, and the data included in sync.
Before sync it summarizes the destination, payload, credential checks, and
configuration to review. Setup finishes after `overmind sync` succeeds with
the mapped capabilities, outstanding verification, and one recommended next
action. Instrumentation, application runs, and server-side evaluator preparation
remain separate work.

Use the project's runner for these commands (`uv run overmind` or
`poetry run overmind`) so an older global installation cannot supply the
skills. An editable SDK install does not refresh skill copies already installed
in a repository. Refresh them with
`uv run overmind skills sync overmind --ide <cursor|claude|opencode|codex>`
(or the equivalent project runner), then start a new conversation so the agent
reads the updated instructions.

## Tracing

Needs `pip install "overmind[tracing]"`. Skip the extra if the app already
pinned OpenTelemetry — use the default install and fan-out (see the telemetry
skill). Wire up once at process start, then annotate the functions you want traced:

```python
import overmind

# Reads OVERMIND_API_KEY or the credential saved by overmind sync. Without a key this logs once
# per process, returns False, and every decorator below becomes a no-op —
# safe to ship.
overmind.init(
    service_name="my-agent",
    capability_id="<capability-uuid>",  # ingest maps traces by this id alone
    capability="Support Triage",  # display label; never resolves anything
    providers="auto",  # instrument every installed provider SDK
)  # (or name them: providers=["openai", "anthropic"])


@overmind.entry_point()  # run root (overmind.unit_kind = "run")
def run(request: dict) -> dict:
    overmind.intent(request["question"])  # what the user asked for
    answer = think(request)
    overmind.deliver(answer)  # terminal deliverable, auto-grounded
    return answer


@overmind.tool(ignore=("session",))  # tool evidence; session never captured
def search(query: str, session) -> list[dict]: ...


@overmind.observe(type="llm", capture="messages")  # full chat evidence
def call_model(messages: list[dict]) -> dict: ...
```

That is the whole integration: no init guards (everything no-ops without a key), no hand-rolled scrubbing (captured payloads redact secret-named keys and base64 blobs automatically, text is kept in full), and no evidence bookkeeping (`deliver()` grounds itself in the environment-provenance spans of the run — pass `grounded_by=[...]` to override). On `KeyboardInterrupt`/cancellation the entry-point span flushes before re-raising, so interrupted runs still land.

Decorators: `entry_point`, `workflow`, `tool`, `retrieval`, and the general `observe` (sync and async). All accept `capture=` (`"auto"` scrubbed args/result, `"none"`, `"messages"`), `ignore=` (argument names never captured), `format_input=` / `format_output=` hooks for custom payload shapes, `provenance=`, `unit=`, and `capability=`. `start_span(...)` is the context-manager companion; `set_tag`, `set_user`, `set_conversation_id`, and `capture_exception` annotate the current span Sentry-style.

The span name may be a callable receiving the call's arguments — for polymorphic dispatchers, where one function executes named actions and each invocation must emit its own tool span (`tool.name` follows the resolved name):

```python
class Tools:
    @overmind.tool(name=lambda self, action, **params: action.name)
    def act(self, action, **params):  # executes navigate / extract / done / ...
        ...
```

Spans declare evidence provenance for the platform's evaluation judges: tool and retrieval spans are tagged `overmind.provenance = "environment"` and LLM spans `"agent"` automatically; pass `provenance=` (`user` / `agent` / `environment` / `harness`) to override. `@entry_point` spans are run roots (`overmind.unit_kind = "run"`) — one per trace: a run declared inside an open trace resolves to `turn`. `unit="turn"` marks an independently scorable decision cycle — each turn becomes one scored task execution. Internal fan-out or iteration spans (parallel sub-queries, retries, loop bodies) must not declare `unit`; handoffs stamp their own `turn` automatically. A `function` span that starts a trace outside any run boundary is an orphan fragment and is not exported by default (`init(export_orphan_spans=True)` overrides).

The wire-level attribute contract is **pinned** in [`docs/tracing-attributes.md`](docs/tracing-attributes.md); nothing there is renamed. The telemetry skill (`skills/overmind/references/telemetry.md`) is the integrator's guide — run vs. turn, deliver placement, handoffs, and the anchor-decoration rule. When traces don't show up: `init(debug=True)` prints the endpoint, identity, enabled instrumentors, and export mode.

Multi-capability agents scope identity with `overmind.capability` — a context manager or decorator that stamps `overmind.capability.id` / `.name` on every span created inside and restores the outer identity on exit (`capability="..."` on any decorator is shorthand for the name-only scope):

```python
with overmind.capability("DOM Element Locator", id="..."):  # id optional
    locate(prompt)  # every span here belongs to the locator capability
```

Entering a different capability mid-trace is a handoff: the first span of the new scope is stamped `overmind.unit_kind = "turn"`, so the platform scores it as a new unit against that capability's evals. Only declared identities are stamped — nothing is auto-created. `overmind.task("behaviour-slug")` optionally pins spans to a declared Behaviour the same way.

Single-capability agents with multiple phases (graph nodes, debate rounds) carve a run into units with `task(..., unit="turn")` — the scope opens one turn span per behaviour per trace, re-entering the same key re-uses it even when a phase's activity is non-contiguous, and the span closes when the run ends:

```python
with overmind.task("investment-debate", unit="turn"):
    ...  # spans here nest under the behaviour's turn span
```

`overmind.run(...)` brackets a whole agent run in one scope — capability identity (args, else `OVERMIND_CAPABILITY_ID` / `OVERMIND_CAPABILITY_NAME`), the entry-point run span, intent, conversation id, tags, error status, and a flush on exit. The yielded handle delivers the terminal payload; call it inside the unit that produced it:

```python
with overmind.run(
    "trading-run", intent=f"Analyze {ticker}", conversation_id=f"{ticker}:{date}"
) as run:
    final_state = app.invoke(state)
    with overmind.task("portfolio-manager", unit="turn"):
        run.deliver(final_state["final_trade_decision"])
```

It is also a decorator (sync or async) for method entry points. Every parameter except `name` accepts a callable receiving the wrapped call's arguments, resolved per invocation, and the run-boundary span carries the function's `code.namespace` / `code.function.name` — one decoration covers both the run bracket and a scan-contract anchor. The return value is not auto-delivered; call `overmind.deliver()` inside the unit that produced it:

```python
class Agent:
    @overmind.run(
        intent=lambda self, *a, **k: self.task,
        conversation_id=lambda self, *a, **k: self.task_id,
    )
    async def run(self): ...
```

### LangChain / LangGraph

`pip install 'overmind[tracing]'`, then `providers=["langchain"]` mounts the OpenInference LangChain instrumentor (covers LangGraph): every chain, LLM and tool invocation gets a span with usable model/token/cost evidence. For the scoring semantics no instrumentor can know, `overmind.integrations.langgraph.bind` maps graph nodes to behaviour turn units — call it on the `StateGraph` after the `add_node` calls, before `compile()`:

```python
from overmind.integrations import langgraph as overmind_langgraph

overmind.init(providers=["openai", "langchain"], capability_id="<capability-uuid>")

workflow = build_state_graph()
overmind_langgraph.bind(
    workflow,
    # Default key per node: slugified node name ("Market Analyst" → "market-analyst").
    # Override where the scanned task map groups nodes differently; None opts a node out.
    behaviours={
        "Bull Researcher": "investment-debate",
        "Bear Researcher": "investment-debate",
        "Msg Clear Market": None,
    },
    deliver="Portfolio Manager",  # optional: this node's completion delivers its return value
)
app = workflow.compile()
```

Each node invocation runs inside `task(key, unit="turn")` (re-entrant phases share one unit) and function-backed nodes carry their `code.namespace` / `code.function.name` identity for contract anchoring.

## MCP and optional plugins

MCP is the complete platform integration. Its tools, resources, initialization
guidance and native workflow prompts work without a plugin. Connect directly
with `overmind init --ide <client>` and `overmind sync`.

The Codex, Cursor and Claude Code plugin manifests package that same connection
and the shared workflow skills. They add installation, branding, workflow guidance and
links to the normal Console. They do not add a separate UI or another API.

| Client      | Install the plugin                                                                      |
| ----------- | --------------------------------------------------------------------------------------- |
| Claude Code | `/plugin marketplace add overmind-core/overmind`, then `/plugin install overmind@overmind` |
| Codex       | `codex plugin marketplace add overmind-core/overmind`, then enable it in `/plugins`        |
| Cursor      | Add `overmind-core/overmind` as a plugin marketplace, then install Overmind              |

The bundled `mcp.json` connects to the hosted API through OAuth account sign-in.
Access continues until revoked, with automatic token refresh. Standalone MCP
also accepts account or project API keys through `X-Api-Key` or
`Authorization: Bearer`; never include a key in a plugin archive. For local or
self-hosted projects, the direct `init`/`sync` setup selects the deployment and
project credential. Keep one active connection for the intended account or
project to avoid duplicate connections.

Call `list_projects`, then read `overmind://project/current?project_id=ID` to
identify the chosen project and obtain its `console_url`. Open that ordinary
Console URL when a visual view is useful;
its browser session remains separate from MCP authentication.

## Skills

Use these from Cursor, Codex, or Claude Code to scaffold agents and operate
Overmind without leaving your coding environment. Skills live at the repo-root
[`skills/`](./skills/) directory so agent installers can pick them up from this
repository (e.g. `npx skills add overmind-core/overmind`).

```bash
overmind skills list --verbose
overmind skills sync overmind
overmind skills sync overmind-observability overmind-datasets --ide codex
overmind init --ide codex
```

`init` prepares each vendor's project-scoped MCP config, leaving any other
configured servers untouched. `sync` installs the final project key:

| `--ide`                  | MCP config                     | Skill install      |
| ------------------------ | ------------------------------ | ------------------ |
| `cursor`                 | `.cursor/mcp.json`             | `.agents/skills`   |
| `claude` / `claude_code` | `~/.claude.json` (local scope) | `.claude/skills`   |
| `opencode`               | `opencode.json`                | `.opencode/skills` |
| `codex`                  | `.codex/config.toml`           | `.agents/skills`   |

Claude Code gets a local-scope server for this directory in `~/.claude.json`
(or `$CLAUDE_CONFIG_DIR/.claude.json`), so the key never enters the repository
and a committed `.mcp.json` is left untouched. Other MCP configs containing the
project key and `.overmind/credentials.toml` are added to the clone-local Git
exclude file and written with owner-only permissions. Sync refuses to put a key
in a tracked config. Codex loads project configuration only for trusted
repositories.

`overmind init` installs the complete skill set. Plugins bundle the same source;
`overmind skills sync` refreshes selected skills in an existing installation.

| Skill                    | What it does                                                       |
| ------------------------ | ------------------------------------------------------------------ |
| `overmind`               | Local setup, repository discovery and workflows across surfaces    |
| `overmind-agent`         | Inspect capabilities, behaviour coverage and repository provenance |
| `overmind-observability` | Investigate traces, failures, latency and instrumentation gaps     |
| `overmind-datasets`      | Prepare, verify, generate and export Data Workshop versions        |
| `overmind-evaluations`   | Author evaluators, prepare runs and compare results                |
| `overmind-optimiser`     | Run prompt/code experiments and model comparisons                  |
| `overmind-training`      | Prepare exact inputs, estimate costs and inspect fine-tuning       |
| `overmind-inference`     | Inspect serving metrics, worker state and live routing             |
| `overmind-integrations`  | Configure provider mappings and verify imported traces             |

## Anonymous usage analytics

The SDK and CLI send anonymous product-analytics events to PostHog so we can
see how the package is adopted. Each CLI process emits one `cli.invoked` event
on exit (`command` = full redacted argv like `overmind skills list`,
`command_path` = nested path like `skills list`, exit code, duration), via
Typer's `call_on_close`. Library calls
emit `sdk_init` / `sdk_client_created` / `sdk_langgraph_bind`. When an API key
is available the SDK identifies the user once (cached under `~/.overmind/`) so
events join the same PostHog person as the Console.

This is **not** agent tracing: no prompts, span payloads, API keys, emails, or
dataset contents are included. Customer OTLP traces still go only to your
Overmind project via `overmind.init()`.

Opt out with any of:

```bash
export OVERMIND_ANALYTICS_ENABLED=false
# or
export DO_NOT_TRACK=1
```

Analytics is also off when `CI` is set in env.

## CLI reference

```text
overmind init [OPTIONS]             Skills, slash commands, MCP; seed overmind.toml
overmind sync [up|down]             Push/pull overmind.toml with the server
overmind chassis [--root PATH]      Record scan provenance and print the AST chassis digest
overmind dataset upload FILE        Upload a local dataset and start a build
                                    (--split PERCENT lands a train and an eval dataset)
overmind dataset export DATASET     Download committed rows as JSONL or CSV
overmind connector add TYPE         Add a tracing connector from env or a TTY prompt
overmind model download-checkpoint DEPLOYMENT
                                    Download an archived fine-tuned checkpoint
overmind optimise [OPTIONS]         SDK loop the /overmind optimise skill drives
overmind skills list [--verbose]    List installed/available skills
overmind skills sync <name>...      Sync one or more skills to the latest version
```

The Console Agent header identifies the repository snapshot behind the capability map. `overmind chassis` records the checkout before scanning; conversion verifies that it stayed unchanged and saves its provenance in `overmind.toml`. `overmind sync` preserves the scan revision and time, with a separate server sync timestamp. Run `/overmind setup` to refresh the map. Live traces may come from other revisions.

Run `overmind <command> --help` for full flag documentation.

## Licence

MIT. See [LICENSE](LICENSE). The Overmind platform that this package talks to is AGPL-3.0, with a commercial licence from Overmind Ltd if you need different terms.
