Metadata-Version: 2.4
Name: sia-foundry
Version: 0.1.4
Summary: SIA Foundry CLI — improve a locally checked-out agent
Author-email: Hexo AI <team@19bits.com>
License-Expression: MIT
Project-URL: Homepage, https://sia.hexo.ai
Project-URL: Documentation, https://sia.hexo.ai
Keywords: sia,foundry,agents,evals,llm,cli
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx>=0.27
Requires-Dist: rich>=13.0
Requires-Dist: PyYAML>=6.0
Requires-Dist: websockets>=13.0
Requires-Dist: harbor>=0.22
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.24; extra == "dev"
Dynamic: license-file

# `sia` — the SIA Foundry CLI

Improve an AI agent you already have checked out. Point `sia` at your repo and
it generates an eval set from your code, runs your real agent against it,
collects traces, diagnoses what is going wrong, and writes patches you review
before they touch your working tree.

```
sia improve --max-cost 5.00
```

## Install

```bash
pip install sia-foundry     # or: uv tool install sia-foundry, pipx install sia-foundry
sia --version
```

One package, and there is no second one to name. The policy that decides what
a server-sent tool call may do to your repo is a single piece of code shared
by both ends of the channel, not two copies that drift, and it ships inside
this wheel as `sia_engine_shared`. The backend installs those same files from
the repo rather than from here, so it never has to pull the CLI in to get
them. The proprietary `sia-engine` package (prompt engineering,
orchestration) is a separate, backend-only distribution — never installed
here.

It also pulls in [Harbor](https://pypi.org/project/harbor/), the sandbox
runner that `sia envs generate` and `sia evals run` shell out to. That is the
bulk of the install; it needs a Docker daemon at runtime, and the rest of the
CLI degrades gracefully without one.

Sign in through your browser — no password in the terminal:

```bash
sia login                   # opens your browser, finishes on its own
sia login --device          # prints a code, for SSH and containers
```

Each login is named after the machine it came from and can be revoked from
**Devices** in the web app.

Requires Python 3.12+. Dependencies: `httpx`, `rich`, `PyYAML`, `websockets`,
`harbor`.

## Getting started

```bash
cd ~/code/my-agent
sia login --foundry https://sia.hexo.ai
sia init
sia status
```

`sia init` registers the project, detects how to run your agent, writes
`.sia/config.toml`, and gitignores the derived artifacts.

Then run the loop:

```bash
sia evals generate    # read the source, synthesize an eval set
sia evals run         # execute your agent, collect traces
sia failures detect   # cluster what went wrong
sia fixes propose     # write unified diffs
sia fixes apply p1    # review the diff, then land it
sia evals run         # confirm the fix
```

`sia` with no arguments opens an interactive shell (`/status`, `/evals`,
`/help`).

## How your agent gets invoked

`sia evals run` executes on **your** machine — it needs your databases,
credentials and local services. Tell it how in `.sia/config.toml`:

```toml
[agent]
kind = "command"                  # command | http | python
cmd  = "python3 -m myagent"
timeout_s = 120
concurrency = 4
```

- **`command`** — a subprocess. Receives `{"input": "...", "case_id": "..."}`
  as JSON on stdin; write `{"output": "..."}` to stdout. Bare text on stdout
  works too, so a script that just prints its answer needs no wrapper.
  `SIA_CASE_ID` is in the environment.
- **`http`** — `url = "http://localhost:8080/chat"`. Receives a POST with
  `{"input", "case_id"}`; accepts `{"output": ...}` or an OpenAI
  chat-completions body.
- **`python`** — `entrypoint = "myagent.main:run"`. Imported from the repo
  root and called with the input string. Sync or async.

A case that crashes or times out is recorded as a failed case, not raised —
a broken agent is usually the thing you are investigating.

## Where the three inputs come from

Every SIA endpoint takes some combination of **source code**, an **eval set**
and **traces**.

**Source code** is your working tree. In a git repo `sia` uses `git ls-files`,
so `.gitignore` is honored exactly; add a `.siaignore` for anything else you
want held back. Tune `[source] include/exclude/max_files/max_bytes` in the
config. Only the selected files are uploaded; every path in the repo is sent as
a manifest so SIA knows what exists.

**The eval set** is a file in your repo: `.sia/evals/default.yaml`. Generate it
with `sia evals generate`, edit it by hand, review it in pull requests. It is
tracked in git on purpose — it is a test asset.

```yaml
version: 1
name: default
cases:
  - id: c1
    input: where is order 5?
    expected_behavior: asks for a customer id before searching
    category: tool-use
    tags: [orders]
```

**Traces** arrive three ways:

1. `sia evals run` emits them itself. Zero setup.
2. `sia traces collect` runs a local OTLP receiver. Point an
   already-instrumented agent at it:
   ```bash
   export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318
   export OTEL_EXPORTER_OTLP_PROTOCOL=http/json
   ```
   (JSON encoding only — that is what keeps protobuf out of the install.)
3. `sia traces pull` fetches what the SIA inference gateway captured. Set
   `[traces] gateway_url` and point your agent's LLM client at it; every call
   is then traced with no instrumentation at all.

The `sia traces` group is unlisted — it works, but it is not in `sia --help`
or the command reference, because ways 1 and 3 need no command at all.

Spans are normalized to the OpenTelemetry GenAI conventions (`gen_ai.*`), so
OpenInference and OTel SDK output map in untranslated.

## Where things run

Reasoning is centralized on the Foundry API; execution stays where your agent
lives.

| Endpoint | Runs | Command |
|---|---|---|
| `generate_evals` | server | `sia evals generate` |
| `generate_evals_from_traces` | server | `sia evals generate --from-traces` |
| `generate_environment` | server | `sia envs generate` |
| `generate_environment_from_traces` | server | `sia envs generate --from-traces` |
| `run_eval` | **local** | `sia evals run` |
| `detect_failures` | server | `sia failures detect` |
| `propose_fix` | server | `sia fixes propose` |
| `apply_fix` | **local** | `sia fixes apply` |

Server-side calls upload the filtered source bundle, the eval set and the
traces. Judging also happens server-side, so the CLI never needs model
credentials.

## Letting SIA into the repo

Your working tree reaches Foundry over a **tool channel**: a WebSocket carrying
one tool call at a time. While it is open, SIA can read your files, edit them,
and run commands here — the same things a coding agent does locally, except the
model runs on the server and your code never leaves the machine.

The running step says which side it is blocked on, and for how long — the model
thinking on the server, or a tool call on this machine:

```
  ⠙ job 22        server     3.0s  waiting on model · 3.0s
  ⠼ job 22        server     5.9s  running Bash cd . && python3 -c "import sys…
```

Both clocks matter: the step's, and this wait's. The same transitions are
timestamped into `.sia/logs/<invocation>.log`, so a run that felt slow can be
read back afterwards and blamed on the right side.

**The commands that need it open it themselves.** `sia improve`,
`sia envs generate`, `sia evals generate`, `sia failures detect` and
`sia fixes propose` all attach when they start and detach when they finish, so
there is no daemon to remember:

```bash
sia improve --max-cost 5.00
```

```
  lending my-agent for this command
✓ Improve run #7 started
```

This is what makes `sia envs generate` able to *prove* a container works rather
than guess: the coding agent runs `harbor run -n 1 --env docker` against your
Docker, with your dependencies.

While one command holds the channel, another started alongside it uses that
session as-is rather than opening a second. The server keeps a single session
per project, so nothing connects behind its back and evicts it.

Every tool call is checked on **your** side before it runs:

| Tool | Policy |
|---|---|
| `Read`, `Glob` | always allowed, confined to the repo |
| `Write`, `Edit` | repo only — `..`, `~`, absolute paths and symlinks out are refused |
| `Bash` | allowlisted (`harbor`, `docker`, `python`, `pytest`, `pip`, read-only `git`, …); anything else asks you first, showing the command |

A refusal goes back to the model as a tool error, so it adapts instead of
failing the run. Nothing about this is configurable from the server. Only
non-allowlisted `Bash` ever prompts; when a command owns the channel itself, the
streaming job log pauses for the question rather than scrolling it away.

The question shows the command in full, then collapses to a single line once
you answer — `✓ allowed Bash: rm -rf traces (5 lines)` — so a multi-line script
does not stay in your scrollback. Both the question and the answer are written
to `.sia/logs/<invocation>.log` in full.

The built-in list is `harbor`, `docker`, `python`, `pytest`, `pip`, read-only
`git`, `ls`, `cat`, `head`, `tail`, `wc`, `jq`, `grep`, `rg`, `find`, `sort`,
`uniq`. Pipelines of those are fine — every stage is checked, so `cat x | sh`
asks about `sh`, not about the pipe.

Being prompted repeatedly means the list is missing something this repo needs.
Name it, rather than turning the gate off:

```toml
[engine]
allow_commands = ["awk", "sqlite3"]
```

Names, not patterns. Allowing a command does not allow what it is chained to:
`awk … && rm -rf ~` still asks, as does anything with `;`, `&&`, backticks,
`$(…)` or a redirect out of the repo.

To skip prompts entirely on a repo you are deliberately handing over, pass `-y`
for one command (`sia improve -y`, `sia fixes propose -y`) or set
`[engine] trust = "full"` for the project.

## The loop

`sia improve` runs detect → fix repeatedly until it stops paying off, instead of
you driving each step. It lends the repo for as long as it runs:

```bash
sia improve --max-cost 5.00
```

It stops on the first of: no failures left, no fix proposed, the cost/accuracy
frontier not moving for two rounds, the budget, or the round limit — and says
which.

**The budget is in dollars, not tokens.** The coding turns are the expensive
part of a round and carry no eval tokens at all, so a token ceiling would bound
the cheap half and let the dear half run free.

Each round is a separate job. If the server restarts mid-round the run is
**parked**, not resumed — re-running a round that already spent money is worse
than stopping and asking — and `sia improve resume` picks it up at the round
boundary. `sia improve cancel` stops it after the current round rather than
tearing down a half-applied patch.

Patches are never applied for you: the loop proposes, and `sia fixes apply` is
still the gate.

## Patch safety

`propose_fix` returns unified diffs plus the SHA of every file each diff was
generated against. `sia fixes apply` refuses if the working tree has moved on
since — the check, not the diff format, is what makes server-side patch
generation safe. Override with `--force` if you know better; undo with
`sia fixes revert <id>`.

Patches are applied all-or-nothing: a hunk that fails on the third file leaves
the first two untouched. A patch that would leave every file byte-identical is
rejected rather than reported as applied.

## Environments (harbor tasks)

`sia envs generate` builds containerized tasks under `.sia/env/<name>/` so eval
cases can run reproducibly: seeded state, pinned dependencies, and a
`tests/verify.py` that exits non-zero on failure. `sia envs validate --build`
docker-builds each one.

## Command reference

```
sia init [path] [--from URL] [--name N] [--foundry URL] [--gateway URL] [--force]
sia login [--foundry URL] [--device]
sia logout [--foundry URL]
sia status

sia evals generate [--from-traces] [-n N] [--name NAME]
sia evals run [--no-harbor] [--concurrency N]
sia evals results
sia evals list | show <case-id>

sia envs generate [--from-traces] [-n N]
sia envs list | validate [--build]

sia failures detect | list | show <failure-id>

sia fixes propose
sia fixes apply [patch-id|all] [-y] [--force]
sia fixes revert <patch-id>
sia fixes list | show <patch-id>

sia                         # interactive shell
```

## Workspace layout

```
.sia/
  config.toml          # committed
  evals/default.yaml   # committed — a review artifact
  env/<task>/          # committed
  state.json           # gitignored
  traces/*.jsonl       # gitignored
  runs/*.json          # gitignored
  failures/*.json      # gitignored
  patches/*.json       # gitignored
  plans/*.json         # gitignored
~/.sia/credentials.json  # API tokens, keyed by Foundry URL, chmod 600
```

Override the credentials path with `SIA_CREDENTIALS`, or supply a token
directly with `SIA_TOKEN`.

## Tests

```bash
cd cli && python -m pytest
```

The end-to-end tests boot the real Foundry API in-process and drive the CLI
through the whole loop against a toy agent, with the model faked and no API
keys required.
