Metadata-Version: 2.5
Name: castia
Version: 0.3.0
Summary: Idiomatic, FastAPI-style Python SDK for Microsoft Foundry hosted agents.
Project-URL: Homepage, https://github.com/sethjuarez/castia
Project-URL: Repository, https://github.com/sethjuarez/castia
Project-URL: Documentation, https://castia.dev
Author: Seth Juarez
License-Expression: MIT
License-File: LICENSE
Keywords: activity,adaptive-cards,agents,fastapi,foundry,microsoft-foundry,sdk,teams
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Application Frameworks
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: azure-ai-projects>=2.3.0
Requires-Dist: azure-identity>=1.19
Requires-Dist: fastapi>=0.115
Requires-Dist: httpx<1,>=0.28.1
Requires-Dist: microsoft-opentelemetry>=1.3.8
Requires-Dist: pyjwt>=2.9
Requires-Dist: uvicorn[standard]>=0.30
Provides-Extra: deploy
Requires-Dist: ruamel-yaml>=0.18; extra == 'deploy'
Provides-Extra: optimize
Requires-Dist: azure-ai-agentserver-optimization>=1.0.0b1; extra == 'optimize'
Requires-Dist: ruamel-yaml>=0.18; extra == 'optimize'
Provides-Extra: test
Requires-Dist: pytest>=8; extra == 'test'
Requires-Dist: ruamel-yaml>=0.18; extra == 'test'
Description-Content-Type: text/markdown

# castia

Idiomatic, FastAPI-style Python SDK for [Microsoft Foundry](https://ai.azure.com)
hosted agents.

`castia` lets a hosted agent speak Foundry's three wire protocols — Activity
(Teams/Bot Framework), OpenAI `responses`, and `invocations` — through
protocol-named decorators, dependency injection (`Depends`), and typed builders
for messages, Adaptive Cards, entities, and invoke envelopes. You decorate a
handler, return a value, and the framework does the rest — the auth chains,
hosting, Activity routing, and telemetry stay out of your file.

## Install

```bash
pip install castia
# or, with uv:
uv add castia
```

Requires Python 3.11+.

## Quickstart

```python
from castia import Agent, Depends, Model, Teams

app = Agent(name="my-agent")

def gpt4o() -> Model:
    return Model("gpt-4o")

@app.message(Teams.direct, Teams.group, Teams.channel_mention)
async def reply(text: str, model: Model = Depends(gpt4o)) -> str:
    return await model.respond(text)

if __name__ == "__main__":
    app.run()
```

Decorate a handler with the surfaces it answers on, return a `str`, and the
framework sends it as the Teams reply. The model is built **once** for the
process and injected via `Depends` — the model choice stays visible in your file
instead of being buried in the framework. (`Model()` with no argument falls back
to `AZURE_AI_MODEL_DEPLOYMENT_NAME`; `get_model` / `use_model("gpt-4o")` are
zero-config conveniences.)

### Composing protocols with routers

Like FastAPI's `include_router`, an `Agent` composes `Router`s so each protocol
can live in its own module:

```python
from castia import Agent
from handlers import activity, responses, invocations

app = Agent(name="my-agent")
app.include(activity.router, responses.router, invocations.router)
```

### Richer replies

Handlers can take a `Message` and reach for typed builders — Adaptive Cards,
suggested actions, citations, mentions, sensitivity labels, live-typing
streamers, and reactions:

```python
from castia import Depends, Message, Model, Reaction, Router, Teams

router = Router()

@router.activity(Teams.direct)
async def reply(text: str, msg: Message, model: Model = Depends(get_model)) -> None:
    await msg.react(Reaction.eyes)
    answer = await model.respond(text)
    await msg.say(answer)
```

## Protocols

`castia` publishes handlers for the protocols in `PUBLISHABLE_PROTOCOLS`:

- **Activity** — Teams / Bot Framework message and invoke turns.
- **`responses`** — the OpenAI `responses` wire shape.
- **`invocations`** — Foundry invoke envelopes (tool execution, agent-to-agent).

## Observability & evaluation

`castia` configures Foundry/Agent 365 telemetry for you when the agent starts.
By default it emits GenAI spans (the `chat {model}` spans the Foundry Traces UI
keys off) but does **not** record the prompt/response **content** onto them.

Recording content is what makes an agent's traces *evaluable* — trace-based
evaluators read the input/output text from the GenAI spans, which is only present
when content recording is enabled. Turn it on deliberately via
`configure_observability`:

```python
from castia.observability import configure_observability

# Records prompt/response text onto GenAI spans so traces can be evaluated.
configure_observability(enable_content_recording=True)
```

Resolution order for each flag is **explicit argument > environment variable >
default**:

| Flag | Argument | Environment variable | Default |
| --- | --- | --- | --- |
| Content recording | `enable_content_recording` | `AZURE_TRACING_GEN_AI_CONTENT_RECORDING_ENABLED` | off |
| GenAI tracing | `enable_genai_tracing` | `AZURE_EXPERIMENTAL_ENABLE_GENAI_TRACING` | on |

Passing nothing preserves the default behavior. Telemetry setup is best-effort:
a failure is logged, never raised, so it can't break startup or a turn.

> **Security caveat:** enabling content recording writes prompt and response
> **text** to Application Insights. Only enable it where storing that content is
> acceptable for your data-handling and privacy requirements.

### Building a scored eval suite

Once your traces are evaluable, `python -m castia eval` wraps the
`azd ai agent eval` extension to synthesize and run a **scored** eval suite — a
generated JSONL dataset plus an auto-generated, weighted **rubric** (a custom
multi-dimension evaluator):

```bash
# Offline gate — validate eval.yaml + rubric files, no Azure, free in CI:
python -m castia eval check

# Synthesize a rubric + dataset from the agent instruction (billable):
python -m castia eval generate --agent my-agent --max-samples 25

# Re-upload locally edited rubric/dataset files as a new version:
python -m castia eval update --evaluator-only

# Submit a scored run against the deployed agent (billable):
python -m castia eval run
```

`check` is a pure, offline referential-integrity gate: it resolves every
evaluator/dataset `local_uri` and validates each rubric dimensions file.
`generate` and `run` submit **billable** Foundry jobs, so both accept
`--dry-run` to print the resolved `azd` command line without submitting
anything. The `azd` wrappers need the build-time extra: `pip install
'castia[deploy]'`.

The rubric dimensions file is a **bare JSON list** where each entry is keyed by
`id` (a stable slug like `correct_outcome`), with an optional
`always_applicable: true` on the catch-all dimension. That cross-SDK shape is
pinned in the monorepo at [`spec/conformance/rubric/`](../../spec/conformance/rubric).

## Optimizer-readiness

The Foundry **Agent Optimizer** searches for a better system prompt (and, when
you declare tools, better tool descriptions) by running candidates against your
eval suite. Making a castia agent *optimizer-ready* is three things: install the
runtime resolver, ship a baseline config, and source your model + instructions
from that config instead of hardcoding them — so the optimizer can swap in a
candidate with **zero handler changes**.

Bind your model dependency with `configured_model()` and thread its resolved
instructions through:

```python
from castia import Depends, Model, Router, configured_model

router = Router()
gpt = configured_model()  # resolves baseline (or the injected candidate) once

@router.responses()
async def reply(text: str, model: Model = Depends(gpt)) -> str:
    return await model.respond(text)   # instructions flow into responses.create
```

`configured_model()` calls `load_agent_config()`, which is **best-effort**: if
the optimizer package isn't installed, resolution fails, or no config is found,
it degrades to environment defaults (`AZURE_AI_MODEL_DEPLOYMENT_NAME`, no
instructions) — the agent runs identically with or without the optimizer.

Ship a baseline under `.agent_configs/baseline/`:

```
.agent_configs/baseline/
  metadata.yaml       # model, instruction_file, (optional) tool_file pointers
  instructions.md     # the system prompt the optimizer tunes
  tools.json          # optional: tool specs the optimizer may reword
```

If your agent declares tools with `app.tools(...)`, keep the baseline
`tools.json` in sync with the code using the build-time reconciler:

```bash
python -m castia optimize          # write/refresh .agent_configs/baseline/tools.json
python -m castia optimize --check   # CI drift gate (exits non-zero, writes nothing)
```

**Config resolution order** is first-wins: `OPTIMIZATION_CONFIG` (inline JSON) →
resolver API (`OPTIMIZATION_CANDIDATE_ID` + `OPTIMIZATION_RESOLVE_ENDPOINT`) →
local `.agent_configs/` → environment defaults. An explicit `config_dir`
argument (or `OPTIMIZATION_LOCAL_DIR`) affects **only** the local source — pass
it anchored to your app root so the baseline resolves the same under
`python app.py` and `python -m castia`. That contract is pinned for every SDK in
[`spec/conformance/optimization/`](../../spec/conformance/optimization).

### Switching to a reasoning (or RFT-tuned) model

The optimizer's model search can land on a **reasoning** model — an o-series or
GPT-5 deployment, or one you mint yourself with reinforcement fine-tuning (RFT).
Those models take a `reasoning.effort` control that plain chat models don't.
`Model` exposes it as `reasoning_effort` (`minimal|low|medium|high`):

```python
o4 = use_model("o4-mini-rft-2025", reasoning_effort="high")
```

An unset effort omits the field entirely, so chat models are called exactly as
before; a bad level raises at construction rather than as a `400` mid-turn. An
operator can also switch a **deployed** agent onto a reasoning model with zero
code by setting `MODEL_REASONING_EFFORT` — an explicit argument still wins.
Because `configured_model()` builds its `Model` through the same path, that env
override flows through to the resolved candidate automatically.

> **Responses-only constraint:** the optimizer accepts only single-protocol
> `responses` agents — submitting a multi-protocol agent (one that also speaks
> activity/invocations) is rejected with a `400` at submission. Project a
> responses-only sibling from the *same* handler code with
> `app.responses_only()`, deploy that as its own service, optimize it, then apply
> the winning `.agent_configs` candidate back to your live agent.

The runtime resolver and the reconciler need the optimizer extra: `pip install
'castia[optimize]'`.

## Reinforcement fine-tuning (RFT)

RFT is the fourth lifecycle step — **build → evaluate → optimize → switch
models**. It trains a *reasoning* model against a **grader** (a reward function)
instead of labeled answers, minting a new fine-tuned deployment that becomes a
candidate in the optimizer's model search. `castia` ships the build-time tooling
to prepare, validate, and (behind one guarded seam) submit an RFT job — the same
three-seam shape as the eval suite: pure builders, offline validators, and one
billable submit seam.

```bash
# Offline gate — validate an RFT dataset (+ grader), no Azure, free in CI:
python -m castia finetune check --dataset train.jsonl --validation val.jsonl --grader grader.json

# Bridge an eval rubric into a score_model grader (offline):
python -m castia finetune grader --rubric rubric.json --model gpt-4o --out grader.json

# Submit a billable RFT job (use --dry-run to print the payload and submit nothing):
python -m castia finetune submit --model o4-mini --dataset train.jsonl \
  --validation val.jsonl --grader grader.json --reasoning-effort high --dry-run
```

A grader is one of `string_check`, `text_similarity`, `score_model`, `python`,
`multi`, or `endpoint` (preview); templates reference two namespaces only —
`{{ sample.output_text }}` and `{{ item.<field> }}`. Datasets are JSONL chat
`messages[]` rows whose **final message role must be `user`**, with extra
top-level keys as the `item.*` ground truth; **both** train and validation
splits are required. `rubric_to_score_model()` bridges an eval rubric straight
into a `score_model` grader — the natural tie between the *evaluate* and
*switch-models* steps.

> ⚠️ **Provisional / doc-derived.** The grader JSON schema, the RFT
> hyperparameter names, and that `fine_tuning.jobs.create` accepts this payload
> are derived from the Foundry RFT how-to and **have not been confirmed against a
> live RFT job**. The builders and validators are fully offline-tested; treat the
> submitted wire shape as provisional until a real submission validates it. The
> language-neutral contract is pinned in `spec/conformance/graders/`.

## Design

`castia` is deliberately import-cheap: `import castia` never pulls in the
instrumented Azure/OpenAI/httpx stacks, so telemetry can be configured before
those libraries load. The heavy imports are deferred into the methods that need
them.

## License

MIT © 2026 Seth Juarez
