Metadata-Version: 2.5
Name: oi-starship-helm
Version: 0.1.0
Summary: Standalone governance SDK wrapping the Microsoft Agent Governance Toolkit (AGT) behind a stable, layer-modular interface.
Author: Orion Innovation
License-Expression: MIT
License-File: LICENSE
License-File: NOTICE
License-File: agent_sre/LICENSE
Keywords: agent-governance-toolkit,ai-agents,governance,kill-switch,langchain,langgraph,policy
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Security
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.10
Requires-Dist: pydantic<3.0,>=2.0
Requires-Dist: pyyaml>=6.0
Provides-Extra: a2a
Requires-Dist: agent-governance-toolkit-core[full]<6,>=5.0; extra == 'a2a'
Provides-Extra: assurance
Requires-Dist: agent-governance-toolkit; extra == 'assurance'
Provides-Extra: dev
Requires-Dist: build>=1.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Provides-Extra: discovery
Requires-Dist: agent-discovery; extra == 'discovery'
Provides-Extra: full
Requires-Dist: agent-discovery; extra == 'full'
Requires-Dist: agent-governance-toolkit; extra == 'full'
Requires-Dist: agent-governance-toolkit-cli; extra == 'full'
Requires-Dist: agent-governance-toolkit-core[full]<6,>=5.0; extra == 'full'
Requires-Dist: agent-marketplace; extra == 'full'
Requires-Dist: agent-rag-governance; extra == 'full'
Requires-Dist: agentmesh-lightning; extra == 'full'
Provides-Extra: identity-trust
Requires-Dist: agent-governance-toolkit-core[full]<6,>=5.0; extra == 'identity-trust'
Provides-Extra: mcp-governance
Requires-Dist: agent-governance-toolkit-core[full]<6,>=5.0; extra == 'mcp-governance'
Provides-Extra: policy
Requires-Dist: agent-governance-toolkit-core[full]<6,>=5.0; extra == 'policy'
Requires-Dist: cedarpy<5,>=4.8; extra == 'policy'
Provides-Extra: rag-governance
Requires-Dist: agent-rag-governance; extra == 'rag-governance'
Provides-Extra: reliability
Provides-Extra: runtime
Requires-Dist: agent-governance-toolkit-core[full]<6,>=5.0; extra == 'runtime'
Provides-Extra: sandbox
Requires-Dist: agent-governance-toolkit-cli; extra == 'sandbox'
Provides-Extra: storage-azure
Requires-Dist: azure-identity>=1.17; extra == 'storage-azure'
Requires-Dist: azure-storage-blob>=12.19; extra == 'storage-azure'
Provides-Extra: storage-cosmos
Requires-Dist: azure-cosmos>=4.5; extra == 'storage-cosmos'
Provides-Extra: storage-gcp
Requires-Dist: google-cloud-storage>=2.14; extra == 'storage-gcp'
Provides-Extra: storage-s3
Requires-Dist: boto3>=1.34; extra == 'storage-s3'
Provides-Extra: supply-chain
Requires-Dist: agent-marketplace; extra == 'supply-chain'
Provides-Extra: training-governance
Requires-Dist: agentmesh-lightning; extra == 'training-governance'
Description-Content-Type: text/markdown

# oi-starship-helm

Install with `pip install oi-starship-helm`; import as `starship_helm`.

A lightweight governance SDK for AI agents: drop it into any agent to
enforce identity, policy, kill-switch and reliability checks on every
action, without that agent having to know it's talking to the Microsoft
Agent Governance Toolkit (AGT) underneath. AGT itself lives behind a
stable, layer-modular interface.

This is the **user-facing** doc — installing the published package, `govern()`'s
call signature, enabling/disabling layers. For the **contributor-facing** doc —
how the SDK is built, the layer pattern, the `core/` facade — see
`CONTRIBUTING.md` in the source distribution. If you're
integrating this SDK into your own backend (especially with an AI coding
assistant's help), start with
`docs/INTEGRATION_GUIDE.md` (also in the source distribution) instead — it's a
step-by-step playbook, including a recipe for running the integration as a
guided conversation with Claude Code.

## Why this exists

Wrapping AGT packages (`agentmesh`, `hypervisor`, `agent_sre`, ...) behind
one governed entry point keeps application code from ever touching AGT
directly. Hand-built per app, that wrapping tends to scatter AGT imports
across many files, so every AGT release becomes a hunt through the codebase.

This SDK packages that pattern once. Every AGT concern lives behind a **port** (a
stable interface) with exactly one **adapter** file that imports the real
AGT package. When AGT changes — a class renames, a package gets deprecated
in favor of another (this already happened three times during this SDK's
own design: `agent-runtime` → `hypervisor`, `agent-mcp-governance` →
deprecated stub with the real logic living in `agentmesh` instead,
`agentmesh.marketplace` → standalone `agent-marketplace`) — **only that one
adapter file needs to change.** Nothing that depends on the port notices.

Every layer is also independently enable/disable-able via config, with no
conditional logic anywhere else: a disabled layer routes to a "noop"
adapter implementing the same port with safe defaults, instead of the real
one.

## Install

Base install plus one extra per layer you want backed by real AGT:

```bash
pip install oi-starship-helm
pip install "oi-starship-helm[identity-trust,policy,runtime,reliability]"  # the 4 core layers
pip install "oi-starship-helm[full]"  # every layer
```

`reliability` is a no-op extra (an empty dependency list) kept only so
the command above still parses — its real dependency, a subset of
`agent_sre`, ships vendored inside this distribution (see
`agent_sre/VENDORED.md`), so the reliability layer's real adapter
works from the base install alone, with no extra to add.

**Requirements:** AGT-backed extras target `agent-governance-toolkit-core`
5.x. A workflow that ships a `policy.rego` also needs the
[OPA CLI](https://www.openpolicyagent.org/docs/latest/#running-opa) on
`PATH` — AGT 5.x has no built-in Rego evaluator, so without it every
governed call for that workflow is denied (`rule="rego_unavailable"`)
rather than enforced partially.

A layer left disabled in config needs no extra installed at all — that's
the point of the noop path. See `pyproject.toml`'s
`[project.optional-dependencies]` for the full extras list, one per layer.

## Quick start

Author one YAML file:

```yaml
# sdk.yaml
storage:
  backend: local
  root_dir: ./telemetry

policies_dir: ./policies/usecases   # <workflow_id>/mesh.yaml + policy.yaml, shared below

identity_trust: {}
policy: {}
runtime: {}
reliability: {}
```

```python
from starship_helm import GovernanceSDK, GovernanceDenied

sdk = GovernanceSDK.from_yaml("sdk.yaml")

def issue_refund(action: dict) -> dict:
    return {"status": "refunded", "order_id": action["order_id"]}

try:
    result = sdk.govern(
        "orders",             # workflow_id
        "support-agent",      # role
        "issue_refund",       # action_type
        {"order_id": "O-1", "amount": 250},
        issue_refund,
        session_id="session-1",
    )
except GovernanceDenied as exc:
    print(f"denied by {exc.rule}: {exc.reason}")
```

`GovernanceSDK.from_yaml(path)` is cached per resolved path, so calling it
at every call site (rather than once at import time) is free after the
first call. `SDKConfig`/`GovernanceSDK(config)` — building the config tree
by hand in Python — still works exactly as before and is the only option
if you need a custom `trust_store`/`backup_store` object or a non-HTTP
mirror target; see `docs/INTEGRATION_GUIDE.md` steps 4 and 6.

Every `govern()` call runs the same 5 steps in order: identity resolve →
kill-switch check → guardrails + policy → trust/SLO update → audit.

`PolicyConfig.policies_dir` above is per-usecase and still required. Every
`govern()` call is also checked against a second, non-usecase-dependent
tier — `compliance`/`harness`/`governance` policy categories, cached
in-memory and read from this SDK's own bundled defaults unless
`PolicyConfig.global_policies_dir` points somewhere else, so no path is
needed to get global policy coverage. See `CONTRIBUTING.md`'s "global
(non-usecase) policy cache" section for how it's cached and reloaded.

## Enabling and disabling layers

Every layer's config has an `enabled` flag. Flipping it is the entire
mechanism — no other code changes:

```python
config.reliability.enabled = False   # SLO/incident recording becomes a no-op
sdk = GovernanceSDK(config)           # govern() still works identically
```

The four layers above (`identity_trust`, `policy`, `runtime`,
`reliability`) default to `enabled=True` and participate in `govern()`.
The other seven are peripheral — disabled by default, invoked on their own
schedule rather than gated per action:

```python
from starship_helm.layers.assurance.config import AssuranceConfig

config.assurance = AssuranceConfig(enabled=True)
sdk = GovernanceSDK(config)

report = sdk.assurance.scan_prompt(system_prompt_text)
if report.is_blocking:
    raise RuntimeError(f"prompt failed assurance scan: grade {report.grade}")

scan = sdk.discovery.scan()          # find agents running in this environment
shadow_agents = sdk.discovery.reconcile()
```

A peripheral layer left disabled either returns a permissive empty result
(assurance, discovery, mcp_governance, rag_governance — these gate nothing
destructive) or raises `LayerDisabledError` when there's no safe
pass-through (sandbox code execution, plugin installation, wrapping an RL
training runner — these perform an action, so "disabled" can't silently
pretend to succeed).

## Storage: where governance state gets exported

The audit trail, identity_trust's trust scores, runtime's kill-switch
state, and (optionally) reliability's SLO/incident snapshot are the SDK's
own runtime state — separate from `policies_dir` (mesh/policy/sre YAML),
which is versioned deployment config a host app ships alongside its code,
not something the SDK exports.

Every audit entry (`sdk.audit.entries`) has a stable, versioned shape
(`starship_helm.audit.AuditEvent`/`SCHEMA_VERSION`): fixed fields
`session_id`, `agent_id` (the role), `use_case` (the workflow id), `action`,
`tool`, `ring`, `result` (`"success"`, `"failure"` or `"denied"`) and
`target_table` (`"session_executions"` for governed actions,
`"kill_switch_events"` for kill-switch trips/disarms), plus a free-form
`attributes` bag for everything else. `starship_helm.audit.to_platform_row(entry)`
and `to_otel_log_record(entry)` translate one stored entry into a flat
database row, or into an OTel-style log record (name/timestamp/
attributes/resource — no `opentelemetry` package required) respectively —
whichever a host's own export pipeline needs. See `CONTRIBUTING.md`'s
"`audit/` — `AuditTrail`, and the canonical event schema it stores"
section for the full field mapping.

All of it is written through one `StoragePort`, selected by a single
`SDKConfig.storage` field:

```python
from starship_helm.core.storage.config import StorageConfig

# Local disk (default) — files under root_dir.
StorageConfig(backend="local", root_dir="./telemetry")

# Export to any HTTP endpoint instead — an Azure Function, an AWS API
# Gateway + Lambda, a GCP Cloud Run service, or a plain REST app. No cloud
# SDK is imported, so the SDK itself stays platform-agnostic: the same
# adapter works unmodified regardless of which cloud is actually listening
# behind endpoint_url.
StorageConfig(
    backend="http",
    endpoint_url="https://my-api.example.com/governance",
    api_key="...",
)

# Process-local, non-persisted — tests and short-lived processes.
StorageConfig(backend="memory")
```

An `http` backend must implement, keyed on `{endpoint_url}/{key}`:

| Method | Behavior |
|---|---|
| `GET` | 200 with the raw bytes previously written, or 404 if the key was never written (or was deleted) |
| `PUT` | store the request body verbatim at that key; any 2xx |
| `DELETE` | remove the key if present; any 2xx or 404 both count as success |

That's the whole contract — an Azure Function backed by Blob Storage, a
Lambda backed by S3 or DynamoDB, or a Cloud Run service backed by GCS or
Firestore all satisfy it identically, so switching where the SDK is
hosted is a matter of pointing `endpoint_url` at that cloud's function,
not a code change here.

Each layer's own `build()` also accepts an optional `storage` argument
directly (`identity_trust.build(config, storage)`), so a layer built
standalone outside a full `GovernanceSDK` still works — it falls back to
an in-memory store if none is given.

Adding a fourth backend (e.g. a native Azure Blob / S3 / GCS SDK adapter
instead of going through HTTP) means adding one file implementing
`StoragePort` (see `core/storage/local_adapter.py` for the shape) and one
branch in `core/storage/__init__.py`'s `build()` — nothing else in the
SDK references a storage backend directly.

## Mirroring state into your own store: the three hooks

`StoragePort` covers "where does the SDK's own state live." Separately,
three generic, best-effort hooks exist for a host app that wants to mirror
that state into a *second*, app-owned store (e.g. a database your own
dashboard queries) without the SDK taking on a dependency on it. Every hook
follows the same contract: called synchronously right after the state it
mirrors is finalized, and any exception it raises is caught and logged —
a mirror failure can never break, deny, or delay a governed call.

| Hook | Set on | Fires | Payload |
|---|---|---|---|
| `audit_on_record` | `SDKConfig` | once per `sdk.audit.record(...)` call | the finished audit entry (`dict`) |
| `ReliabilityConfig.on_error_budget` | `ReliabilityConfig` | once per SLO, after every `record_outcome`/`emit_*_signal` | `{slo_id, agent_id, use_case, budget_total, budget_consumed, period_start, period_end, updated_at}` |
| `IdentityTrustConfig.trust_store` | `IdentityTrustConfig` | on every trust resolve/update | not a callback — a full `TrustStorePort` you implement (`load`/`seed`/`record_event`) |

If your mirror target for the first two rows is a plain HTTP endpoint,
`GovernanceSDK.from_yaml`'s `http:` shorthand (see
`docs/INTEGRATION_GUIDE.md` step 4) builds the callable for you — no Python
needed. `trust_store` (and anything that isn't a plain HTTP POST) still
means hand-building `SDKConfig` as below.

```python
config = SDKConfig(
    audit_on_record=lambda entry: my_db.insert("audit_events", entry),
    reliability=ReliabilityConfig(
        policies_dir="./policies/usecases",
        on_error_budget=lambda row: my_db.upsert("error_budgets", row, key="slo_id"),
    ),
    identity_trust=IdentityTrustConfig(
        policies_dir="./policies/usecases",
        trust_store=MyTrustStore(),  # implements load/seed/record_event -- see
                                      # starship_helm/layers/identity_trust/trust_store.py
    ),
    ...
)
```

Leave any of these unset and behavior is unchanged from before the field
existed: audit entries still get written (just not mirrored), reliability
still persists a snapshot if `persist_snapshot=True`, and trust scores
persist through the default `StorageBackedTrustStore` (current-value-only,
via the shared `StoragePort`).

## All layers

| Layer | Config class | Wraps | In `govern()`? |
|---|---|---|---|
| `identity_trust` | `IdentityTrustConfig` | `agentmesh.client.AgentMeshClient` | Yes |
| `policy` | `PolicyConfig` | `agentmesh.governance.govern()` + `Policy` + approval flow | Yes |
| `runtime` | `RuntimeConfig` | `hypervisor.security.kill_switch.KillSwitch` | Yes |
| `reliability` | `ReliabilityConfig` | `agent_sre` SLO + incidents | Yes |
| `sandbox` | `SandboxConfig` | `agent_sandbox.docker_provider` | No — standalone |
| `assurance` | `AssuranceConfig` | `agent_compliance` verify/prompt-defense/lint | No — standalone |
| `supply_chain` | `SupplyChainConfig` | `agent_marketplace` installer/signing/trust tiers | No — standalone |
| `rag_governance` | `RagGovernanceConfig` | `agent_rag_governance.RAGGovernor` | No — standalone |
| `mcp_governance` | `McpGovernanceConfig` | `agentmesh.services.behavior_monitor.AgentBehaviorMonitor` | No — standalone |
| `training_governance` | `TrainingGovernanceConfig` | `agent_lightning_gov` runner/reward | No — standalone |
| `discovery` | `DiscoveryConfig` | `agent_discovery` scan/reconcile/risk | No — standalone |
| `a2a` | `A2AConfig` | `agentmesh.integrations.a2a` (`A2AAgentCard` + `A2ATrustProvider`) + `agentmesh.identity.AgentIdentity` | No — standalone |

Every layer, real AGT package installed or not, is reachable on the SDK
instance: `sdk.identity`, `sdk.policy`, `sdk.runtime`, `sdk.reliability`,
`sdk.sandbox`, `sdk.assurance`, `sdk.supply_chain`, `sdk.rag_governance`,
`sdk.mcp_governance`, `sdk.training_governance`, `sdk.discovery`, `sdk.a2a`.

### `a2a`: Agent2Agent protocol support

Mints this agent's signed A2A card, verifies a peer before delegating a
task to it, and creates that task:

```python
config.a2a = A2AConfig(enabled=True, sponsor="ops@example.com")
sdk = GovernanceSDK(config)

card = sdk.a2a.issue_agent_card(
    "orders", "support-agent", url="https://support-agent.example.com",
    capabilities=["issue_refund"],
)
if sdk.a2a.verify_peer("orders", "support-agent", peer_did="did:mesh:..."):
    task = sdk.a2a.create_task("orders", "support-agent", peer_did="did:mesh:...", message={"order_id": "O-1"})
```

Peer verification here covers the trust-bridge/cache path only — it does
not fetch and verify a remote peer's AI Card from its
`.well-known/ai-card.json` endpoint (what the A2A wire format expects
instead of an embedded card). Wiring in a real
`agentmesh.trust.bridge.TrustBridge` is the natural next step; not
implemented yet. `AgentIdentity`'s DID also isn't stable across process
restarts — only its public key round-trips (`to_jwk`/`from_jwk`); the
private signing key lives in an AGT keystore this layer doesn't manage.

## Adding support for a new AGT version

1. Check `starship_helm/compat/versions.py`
   — `ADAPTER_TARGETS` names which AGT package and import root each
   layer's `agt_adapter.py` targets. Start there to find the right file.
2. Edit that layer's `agt_adapter.py` to match the new API. The `port.py`
   in the same folder should not need to change — if it does, that's a
   breaking change for every caller, not routine AGT-version drift.
3. Update the corresponding entry in `ADAPTER_TARGETS`.

Nothing outside that one layer folder references the AGT package directly,
so nothing else needs touching.

## Authoring a new layer

Every layer is a folder under `starship_helm/layers/<name>/`
with four files, following `layers/identity_trust/` as the reference
example:

- `config.py` — a `LayerConfig` subclass (from `layers/base.py`) with
  `enabled: bool` plus whatever the real adapter needs.
- `port.py` — a `Protocol` defining the stable interface. Never imports
  AGT.
- `agt_adapter.py` — the only file that imports the real AGT package.
  Translates AGT exceptions/types into `core.errors` / `core.types`
  shapes.
- `noop_adapter.py` — implements the same port with permissive defaults
  (or raises `LayerDisabledError` if there's no safe default — see
  `layers/sandbox/noop_adapter.py` for that case).
- `__init__.py` — exposes `build(config) -> Port`, lazily importing
  `agt_adapter` only when `config.enabled` is true (so importing the
  package never requires the AGT dependency to be installed).

Register the new layer's config in `core/config.py`'s `SDKConfig`, wire it
into `core/sdk.py`'s `GovernanceSDK.__init__` (and `govern()` if it's a
core, per-action layer), and add it to `core/registry.py`'s
`LAYER_MODULES` so the generic test suite picks it up automatically.

## Running the tests

```bash
pip install -e ".[dev]"
pytest tests
```

The suite passes with **zero AGT packages installed** — every test either
exercises a noop adapter directly, or is skipped (not failed) via
`pytest.importorskip` when the real AGT package it needs isn't present.
Install any layer's extra (e.g. `pip install -e ".[identity-trust]"`) to
also exercise that layer's real adapter import. `reliability` is the one
exception: its real adapter (`agent_sre`, vendored — see
`agent_sre/VENDORED.md`) is always importable, extra or not, so its
`test_real_agt_smoke.py` case never skips.

## Building a wheel

```bash
pip install build
python -m build --wheel
```

This directory (`pyproject.toml`'s own location) is the build root. Its
`[tool.hatch.build.targets.wheel]`'s `only-include = ["starship_helm", "agent_sre"]`
ships exactly those two top-level packages — `starship_helm` and the vendored
`agent_sre` — and nothing else from this directory (`tests/`, docs,
`pyproject.toml` itself) ends up in the wheel. A new top-level package added
later needs a one-line addition to `only-include`, not a new remap.
