Metadata-Version: 2.4
Name: sinan-agentic-core
Version: 0.16.1
Summary: State-of-the-art framework for building AI agents with OpenAI Agents SDK
Author-email: Thomas Pernet-Coudrier <t.pernetcoudrier@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/thomaspernet/sinan-agentic
Project-URL: Documentation, https://github.com/thomaspernet/sinan-agentic#readme
Project-URL: Repository, https://github.com/thomaspernet/sinan-agentic
Project-URL: Issues, https://github.com/thomaspernet/sinan-agentic/issues
Keywords: ai,agents,llm,openai,multi-agent,orchestration
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: openai-agents<0.23,>=0.21.1
Requires-Dist: openai<4,>=3.0.0
Requires-Dist: pydantic<3,>=2.12.2
Requires-Dist: pyyaml<7,>=6.0
Provides-Extra: mcp
Requires-Dist: mcp<2,>=1.19.0; extra == "mcp"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.21; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: black>=23.0; extra == "dev"
Requires-Dist: ruff>=0.1.0; extra == "dev"
Requires-Dist: mypy>=1.0; extra == "dev"
Requires-Dist: mcp<2,>=1.19.0; extra == "dev"
Requires-Dist: types-PyYAML>=6.0; extra == "dev"
Requires-Dist: httpx2<3,>=2.7.0; extra == "dev"
Dynamic: license-file

# Sinan (司南)

<img src="https://upload.wikimedia.org/wikipedia/commons/thumb/5/53/Model_Si_Nan_of_Han_Dynasty.jpg/250px-Model_Si_Nan_of_Han_Dynasty.jpg" alt="Sinan - a Han dynasty south-pointing spoon on a bronze plate" width="220" align="right" />

> **sinan (司南)** - the earliest known compass. A lodestone carved into a spoon, resting on a bronze plate inscribed with the 24 directions. Han dynasty China, ~2nd century BCE.
>
> Like its namesake, this framework is the instrument that helps you align yourself with our agentic logic - the plate holds the field of tools, knowledge, and rules; the spoon is the agent that always knows which way to turn.

A framework for building AI agents using the OpenAI Agents SDK. Fork this repository to quickly create agent-based applications.

## Features

- **Declarative Agents** - Define agents as data, not code
- **Registry Pattern** - Central registries for agents, tools, and guardrails
- **Agent Factory** - `create_agent_from_registry()` builds an `Agent` in one call
- **Chat Service** - `chat()`, `chat_with_hooks()`, `chat_streamed()` for API endpoints
- **Token Usage Tracking** - Every chat/run returns token counts (input, output, cached, reasoning)
- **RunHooks** - Track tool calls in real time via `StreamingRunHooks`
- **Session** - In-memory or SQLite-backed conversation history
- **Dynamic Context** - Pass runtime data to instructions and tools
- **InstructionBuilder** - Base class for section-based instruction assembly (persona, context, steps, rules, output)
- **Output Models** - `ToolOutput` and `ChatResponse` dataclasses
- **Agent Catalog** - Load agent definitions (model, description, tools) from a YAML file
- **Tool Catalog** - Load tool metadata (description, category, recovery hints) from a YAML file
- **Knowledge Store** - Inject domain knowledge from YAML files into agent system prompts
- **Structured Agent-as-Tool** - Typed input schemas and structured error handling for sub-agent calls
- **Model Retry Policies** - Declare `model_retry` on an agent and transient model-API failures retry inside the SDK instead of failing the run
- **Model Call Timeout** - Declare `model_timeout` on an agent and no single model-call attempt hangs past the declared number of seconds
- **Tool Output Trimming** - Declare `tool_output_trim` on an agent and oversized tool outputs from older turns shrink before each model call, so the run overflows less often
- **Capabilities** - Pluggable agent behaviors (turn budgets, error recovery, tool tracing, custom hooks) - write a `Capability` subclass and attach it to an `AgentDefinition`. See [`documentation/project/capabilities.md`](documentation/project/capabilities.md).
- **MCP Server** - Expose registered tools as an MCP server (stdio or HTTP) with zero description duplication - all metadata comes from `tools.yaml`

## Installation

```bash
pip install sinan-agentic-core
export OPENAI_API_KEY="your-key"
```

## Usage

### 1. Register a tool and an agent

```python
from agents import function_tool
from sinan_agentic_core import register_tool, AgentDefinition, register_agent

@register_tool(name="get_weather")
@function_tool
async def get_weather(ctx, city: str) -> dict:
    return {"temperature": 72, "conditions": "sunny"}

register_agent(AgentDefinition(
    name="weather_assistant",
    description="Helps with weather queries",
    instructions="You help users get weather information.",
    tools=["get_weather"],
    model="gpt-4o-mini",
))
```

### 2. Run the agent

```python
from agents import Runner
from sinan_agentic_core import create_agent_from_registry

agent = create_agent_from_registry("weather_assistant")
result = await Runner.run(agent, "What's the weather in Paris?")
print(result.final_output)
```

## Chat Service

Ready-to-use functions for API endpoints. All accept an optional `context=` parameter that gets forwarded to `Runner.run()`, making it available to dynamic instructions and tools via `RunContextWrapper`.

```python
from sinan_agentic_core import chat, chat_with_hooks, chat_streamed, AgentSession

session = AgentSession(session_id="user-123")

# --- Non-streaming ---
result = await chat("What's the weather?", "weather_assistant", session)
# {"success": True, "response": "...", "session_id": "...", "tools_called": [...], "usage": {...}}

# --- Streaming with tool notifications (RunHooks) ---
async for event in chat_with_hooks("What's the weather?", "weather_assistant", session,
                                    tool_friendly_names={"get_weather": "Checking weather"}):
    if event["event"] == "tool_start":
        print(f"Tool: {event['data']['friendly_name']}")
    elif event["event"] == "answer":
        print(event["data"]["response"])

# --- Token-level streaming ---
async for event in chat_streamed("What's the weather?", "weather_assistant", session):
    if event["event"] == "text_delta":
        print(event["data"]["delta"], end="", flush=True)
    elif event["event"] == "answer":
        print(f"\nTools used: {event['data']['tools_called']}")
```

### Streaming the model's reasoning

A reasoning model can narrate its own thinking, and the token-streaming paths —
`chat_streamed()` and `BaseAgentRunner.execute(streaming=True)` — forward that
narration as it is written. Ask for it on the agent's `model_settings`; nothing
else is needed, and an agent that does not ask for it behaves exactly as before.

```python
from agents import Agent, ModelSettings
from openai.types.shared import Reasoning

agent = Agent(
    name="analyst",
    model="gpt-5",
    model_settings=ModelSettings(reasoning=Reasoning(effort="medium", summary="auto")),
)

async for event in chat_streamed("Why did revenue drop?", agent=agent, session=session):
    if event["event"] == "reasoning_delta":
        print(event["data"]["delta"], end="", flush=True)   # a thought, as it is written
    elif event["event"] == "reasoning_part_done":
        print()                                             # that thought ended
    elif event["event"] == "text_delta":
        print(event["data"]["delta"], end="", flush=True)
```

| Event | Payload | Meaning |
|---|---|---|
| `reasoning_part_added` | `{"index": 0}` | a new thought starts |
| `reasoning_delta` | `{"delta": "...", "index": 0}` | the next chunk of thought `index` |
| `reasoning_part_done` | `{"index": 0, "text": "..."}` | thought `index` is finished |
| `reasoning` | `{"summary": ["...", "..."]}` | every finished thought, repeated once at the end |

`index` is what separates one thought from the next — a step log renders each as
its own paragraph. The terminal `reasoning` event repeats them all, the same
guarantee `answer` gives for the response text, so a consumer that dropped a
fragment still ends up with exactly what was said.

The model decides how much to narrate, and how much arrives varies with the
model and its `effort` setting. A turn it had nothing to say about produces no
reasoning events at all — normal, rather than a dropped event, so a consumer
must render that case. What arrives is the model's own summary of its thinking,
not its raw internal tokens.

The other two paths report the same summaries, but whole rather than as they
are written, since neither streams the model's tokens: `chat()` returns them as
a `reasoning` list beside the response, and `chat_with_hooks()` yields one
`reasoning` event just before the answer. `BaseAgentRunner.last_reasoning`
holds them after a non-streamed run.

### Reporting a failed run

None of the three raises on a failed run — the failure comes back as data. Alongside the rendered message, each reports the *kind* of failure, classified by exception type through `classify_run_error()` (see [Structured error handling](#structured-error-handling)). `AgentOrchestrator.run_workflow()` reports it the same way.

```python
from sinan_agentic_core import RunErrorKind

result = await chat("What's the weather?", "weather_assistant", session)
# {"success": False, "error": "Max turns (10) exceeded", "error_kind": "max_turns", "session_id": "..."}

async for event in chat_streamed("What's the weather?", "weather_assistant", session):
    if event["event"] == "error":
        # {"error": "Model refused to produce output: I can't help with that.",
        #  "error_kind": "model_refusal"}
        if event["data"]["error_kind"] == RunErrorKind.MODEL_REFUSAL:
            ...
```

`error_kind` is a `RunErrorKind` value — `max_turns`, `context_overflow`, `model_refusal`, `model_behavior`, `input_guardrail_tripwire`, `output_guardrail_tripwire`, or `unknown` — carried as a plain string so the payload stays JSON-serializable. `RunErrorKind` is a `str` enum, so a comparison against either the member or its value works. Branch on it to decide whether a retry is worth attempting; an API layer that matched the message text instead would break the next time upstream rewords it.

When a guardrail stops the run, the report carries a `guardrail` entry as well. The SDK's message names the guardrail's *class*, which is `InputGuardrail` for every input guardrail an agent declares, so the message alone never says which check rejected the request. The entry names the guardrail as it was registered — the name `agents.yaml` writes — and lists every guardrail that finished before the run stopped.

```python
result = await chat("My card number is ...", "weather_assistant", session)
# {"success": False,
#  "error": "Guardrail InputGuardrail triggered tripwire",
#  "error_kind": "input_guardrail_tripwire",
#  "guardrail": {"name": "blocks_pii",
#                "results": [{"name": "off_topic", "tripwire_triggered": False},
#                            {"name": "blocks_pii", "tripwire_triggered": True}]},
#  "session_id": "..."}
```

All three chat functions and `run_workflow()` report the same set. Neither tripwire kind is retryable: a guardrail is a declared check saying no, so a second call that reaches a different answer has defeated it rather than recovered from a limit.

## Token Usage Tracking

Every path that runs a model reports the same record — the chat functions return it, `BaseAgentRunner.run_agent()` returns it, and `execute()` leaves it on `runner.last_usage`. It carries what the provider billed for every model call of the run, the rescue call `execute(fallback_on_overflow=True)` makes included.

```python
result = await chat("What's the weather?", "weather_assistant", session)
print(result["usage"])
# {
#     "requests": 2,
#     "input_tokens": 1500,
#     "output_tokens": 350,
#     "total_tokens": 1850,
#     "input_tokens_details": {"cached_tokens": 200},
#     "output_tokens_details": {"reasoning_tokens": 0},
# }

# Streaming functions include usage in the final "answer" event:
async for event in chat_streamed("Hello", "my_agent", session):
    if event["event"] == "answer":
        print(f"Tokens used: {event['data']['usage']['total_tokens']}")
```

`cached_tokens` is the count the provider served from its prompt cache, summed over the run's calls. A zero therefore means the cache was genuinely cold, not that the number went uncounted — which is what makes it usable as evidence when tuning a prompt for cache hits.

### Pinning the prompt cache shard

OpenAI prompt caching is prefix-based, and the provider shards its cache by key: two calls that share a leading span reuse it only when they route to the same shard. The SDK derives a key from the conversation, session, or group id, but only for an official OpenAI client — an Azure deployment gets none.

Pass `prompt_cache_key` to name the shard yourself. It reaches every model call of the run, the overflow rescue call included, and the three chat functions accept it for the same reason.

```python
output = await runner.execute(
    "weather_assistant", context, session,
    prompt_cache_key=f"tenant:{tenant_id}",
)

result = await chat("What's the weather?", "weather_assistant", session,
                    prompt_cache_key=f"tenant:{tenant_id}")
```

A key you already placed in the agent's own `ModelSettings` — under `extra_args` or `extra_body` — wins: the framework leaves it alone rather than sending a second, conflicting value.

## Dynamic Context

Pass runtime data to agent instructions and tools via `context=`.

```python
from dataclasses import dataclass
from agents import Agent, Runner, RunContextWrapper

@dataclass
class UserContext:
    user_name: str
    language: str = "English"

def dynamic_instructions(ctx: RunContextWrapper[UserContext], agent: Agent) -> str:
    return f"You are a helpful assistant for {ctx.context.user_name}. Respond in {ctx.context.language}."

agent = Agent[UserContext](name="assistant", instructions=dynamic_instructions, model="gpt-4o-mini")

result = await Runner.run(agent, "Hello!", context=UserContext(user_name="Thomas"))

# Works the same with chat functions:
result = await chat("Hello!", "my_agent", session, context=UserContext(user_name="Thomas"))
```

## InstructionBuilder

Build agent system instructions with a consistent section pattern instead of ad-hoc string concatenation. Subclass `InstructionBuilder`, override the sections you need, and skip the rest.

```python
from sinan_agentic_core import InstructionBuilder, AgentDefinition, register_agent

class MyAgentBuilder(InstructionBuilder):
    def __init__(self, context, agent_def):
        super().__init__(context, agent_def)
        self._config = self._ctx_attr("my_config", {})

    def persona(self):
        return self.format_persona("data analyst", self._config.get("persona"))

    def steps(self):
        return self.format_steps(["Load the dataset.", "Analyze trends.", "Produce output."])

    def rules(self):
        return self.format_rules(["Do not fabricate data.", "Cite sources."])

    def output_format(self):
        return 'Output: {"analysis": "...", "confidence": 0.95}'

register_agent(AgentDefinition(
    name="analyst",
    description="Analyzes data",
    instructions=MyAgentBuilder.callable(),  # (context, agent_def) -> str
    tools=["read_data"],
))
```

`build()` assembles sections in order (persona → domain_knowledge → context → steps → rules → output), skips any that return `None`, and joins with double newlines. Override `sections()` to reorder, or `extra_sections()` to append additional `(header, body)` blocks.

For agents with fundamentally different instruction paths (e.g., surface vs deep extraction), use a shared private base with concrete subclasses and a dispatcher function - see `sinan_agentic_core/instructions/builder.py` docstring for details.

## Agent Catalog (YAML-driven agent config)

Keep static agent config (model, description, tools) in a YAML file instead of scattered across Python files. Dynamic parts (instructions, output_dataclass, hosted_tools) stay in Python.

### Basic usage

```yaml
# agents.yaml
agents:
  weather_assistant:
    model: gpt-4o-mini
    description: Helps with weather queries
    tools:
      - get_weather
      - get_forecast
```

```python
from sinan_agentic_core import load_agent_catalog, AgentDefinition, register_agent

catalog = load_agent_catalog("agents.yaml")
cfg = catalog.get("weather_assistant")
# cfg.model -> "gpt-4o-mini"
# cfg.tools -> ["get_weather", "get_forecast"]

register_agent(AgentDefinition(
    name="weather_assistant",
    model=cfg.model,
    description=cfg.description,
    tools=cfg.tools,
    instructions=build_weather_instructions,  # dynamic - stays in Python
))
```

### Tool groups

Define reusable tool sets once, reference them with `group:`.

```yaml
tool_groups:
  graph_navigation:
    - discover
    - search
    - explore
    - read
  project_knowledge:
    - get_rules
    - get_skills

agents:
  research_agent:
    model: gpt-4o
    description: Deep research agent
    tools:
      - think
      - group: graph_navigation      # expands to 4 tools
      - group: project_knowledge     # expands to 2 tools
```

### Conditional tools

Include a tool only when a config flag is truthy. Pass your config object to `catalog.get()` -- the `when:` path is resolved via `getattr()`.

```yaml
agents:
  chatbot:
    model: gpt-4o
    description: Main assistant
    tools:
      - think
      - search
      - tool: web_search
        when: features.web_search_enabled   # dot-path into your config
```

```python
cfg = catalog.get("chatbot", config=my_config)
# If my_config.features.web_search_enabled is True:
#   cfg.tools -> ["think", "search", "web_search"]
# If False or missing:
#   cfg.tools -> ["think", "search"]
```

### Agent-level conditions

Control entire agent registration with a `when:` clause. Use `catalog.is_enabled()` to check.

```yaml
agents:
  web_search_agent:
    model: gpt-4o-mini
    when: features.web_search_enabled
    description: Search the internet
    tools: []
```

```python
if catalog.is_enabled("web_search_agent", config=my_config):
    cfg = catalog.get("web_search_agent", config=my_config)
    register_agent(AgentDefinition(
        name="web_search_agent",
        model=cfg.model,
        description=cfg.description,
        tools=cfg.tools,
        hosted_tools=[get_web_search_tool],  # SDK-specific - stays in Python
    ))
```

### API reference

| Method | Returns | Description |
|--------|---------|-------------|
| `catalog.get(name, config=None)` | `AgentYamlEntry` | Resolved entry (groups expanded, conditions evaluated) |
| `catalog.is_enabled(name, config=None)` | `bool` | Check agent-level `when` condition |
| `catalog.list_agents()` | `list[str]` | All agent names in the catalog |

## Guardrails

A guardrail's `category` decides where it is wired into the SDK. Register the guardrail once in Python, then list it by name on any agent — in `agents.yaml` or directly on an `AgentDefinition`.

| Category | SDK slot | Runs |
|----------|----------|------|
| `input` | `Agent(input_guardrails=...)` | Before the agent starts, on the run input |
| `output` | `Agent(output_guardrails=...)` | After the agent finishes, on the final output |
| `tool_input` | `FunctionTool(tool_input_guardrails=...)` | Before a tool executes, on the tool call arguments |

`tool_input` guardrails are proactive: a rejecting guardrail returns its message as the tool output and the tool never runs. Agents run through `BaseAgentRunner` or the chat functions also run their tool-input guardrails *before* the SDK emits a pending human-approval interruption, so a bad call is stopped without bothering a reviewer. That ordering is a run-level setting (`RunConfig.tool_execution`), so driving `Runner` yourself means passing it in — `build_run_config()` reads it off the agent:

```python
from agents import Runner
from sinan_agentic_core import build_run_config, create_agent_from_registry

agent = create_agent_from_registry("my_agent")
result = await Runner.run(agent, "Hello!", run_config=build_run_config(agent))
```

The same call also carries a declared [tool-output trim policy](#tool-output-trimming), the other run-level setting an agent cannot ride in on. It returns `None` when the agent needs neither, which `Runner.run()` accepts as "use the defaults".

```python
from agents import (
    GuardrailFunctionOutput,
    ToolGuardrailFunctionOutput,
    input_guardrail,
    tool_input_guardrail,
)
from sinan_agentic_core import GuardrailCategory, register_guardrail

@register_guardrail(
    name="block_empty_query",
    description="Reject empty user queries",
    category=GuardrailCategory.INPUT,
)
@input_guardrail
async def block_empty_query(ctx, agent, agent_input):
    return GuardrailFunctionOutput(
        output_info=None,
        tripwire_triggered=not str(agent_input).strip(),
    )

@register_guardrail(
    name="block_destructive_cypher",
    description="Reject write clauses in read-only queries",
    category=GuardrailCategory.TOOL_INPUT,
)
@tool_input_guardrail
def block_destructive_cypher(data):
    if "DELETE" in data.context.tool_arguments.upper():
        return ToolGuardrailFunctionOutput.reject_content(
            message="Destructive Cypher is not allowed. Use a read-only MATCH query.",
        )
    return ToolGuardrailFunctionOutput.allow()
```

```yaml
# agents.yaml
agents:
  graph_agent:
    model: reasoning
    description: Queries the knowledge graph
    tools: [run_cypher]
    guardrails:
      - block_empty_query
      - block_destructive_cypher
```

```python
cfg = catalog.get("graph_agent")
register_agent(AgentDefinition(
    name="graph_agent",
    model=cfg.model,
    description=cfg.description,
    instructions="...",
    tools=cfg.tools,
    guardrails=cfg.guardrails,
))
```

Tool-input guardrails attach to every local function tool the agent resolves. Registry tools are copied first, so one agent's guardrails never leak into another agent using the same tool. Hosted tools (web search, file search) are left untouched — the SDK runs tool-input guardrails for local function tools only.

Both agent-building paths wire declared guardrails the same way: `BaseAgentRunner.create_agent()` and `create_agent_from_registry()`.

The registered `name` is the guardrail's only identifier. Registration stamps it onto the SDK guardrail object, which otherwise falls back to the decorated function's `__name__` — so the tripwire report, the trace span, and the catalog all name the guardrail the way `agents.yaml` does. See [Reporting a failed run](#reporting-a-failed-run) for what a tripped guardrail reports back.

## Tool Catalog (YAML-driven tool metadata)

Keep static tool metadata (description, category, parameters, recovery hints) in a YAML file instead of repeating it in every `@register_tool()` decorator. The Python decorator becomes minimal - just a name linking the function to its YAML entry.

### Basic usage

```yaml
# tools.yaml
tools:
  search_database:
    description: >-
      Search the database for records matching a query.
      Supports semantic and keyword search modes.
    category: search
    parameters_description: "query (str): Search text. search_type (str): 'semantic' or 'keyword'."
    returns_description: "JSON with matching records"
    recovery_hint: "Requires a non-empty query. Prefer search_type='semantic' for natural language."

  read_record:
    description: Read a single record by UUID
    category: search
    parameters_description: "uuid (str): Record UUID"
    returns_description: "JSON with record data"
    recovery_hint: "Verify the UUID is valid. Use search_database to find the correct UUID."
```

```python
from agents import function_tool
from sinan_agentic_core import register_tool, load_tool_catalog, get_tool_registry

# Decorator is minimal - just links function to name
@register_tool(name="search_database")
@function_tool
async def search_database(ctx, query: str, search_type: str = "keyword") -> str:
    ...

@register_tool(name="read_record")
@function_tool
async def read_record(ctx, uuid: str) -> str:
    ...

# At startup: load YAML and enrich the registry
catalog = load_tool_catalog("tools.yaml")
catalog.enrich_registry(get_tool_registry())
```

Those five fields plus the optional `mcp` block are the whole of a tool entry, and an unrecognized key there is rejected when the catalog resolves the tool — a misspelled `recovery_hints` fails at startup rather than leaving the tool without a hint.

### How the merge works

Registration happens in two phases:

1. **Import time** - `@register_tool(name="search_database")` registers the function with an empty `ToolDefinition` (name + function only)
2. **Startup** - `catalog.enrich_registry(registry)` patches each `ToolDefinition` with metadata from the YAML file

YAML values always win over decorator values. Empty YAML fields do not overwrite existing decorator values. This makes it backward compatible - decorators with full metadata still work without a YAML file.

### Backward compatibility

The decorator still accepts all metadata fields. These two approaches are equivalent:

```python
# Approach 1: YAML-driven (preferred for large projects)
@register_tool(name="search_database")

# Approach 2: Decorator-driven (still works, no YAML needed)
@register_tool(
    name="search_database",
    description="Search the database",
    category="search",
    parameters_description="query (str): Search text",
    returns_description="JSON with results",
)
```

If both are provided, YAML wins for non-empty fields.

### API reference

| Method | Returns | Description |
|--------|---------|-------------|
| `catalog.get(name)` | `ToolYamlEntry` | Resolved entry for a single tool |
| `catalog.list_tools()` | `list[str]` | All tool names in the catalog |
| `catalog.enrich_registry(registry)` | `None` | Patch registry ToolDefinitions with YAML metadata |

## MCP Server (optional)

Expose your registered tools as an [MCP](https://modelcontextprotocol.io/) server so any MCP-compatible client (Claude Desktop, Claude Code, VS Code, Cursor, ChatGPT Desktop) can call them. Tool descriptions, parameter schemas, and annotations all come from `tools.yaml` - no duplication.

Requires the `mcp` extra:

```bash
pip install 'sinan-agentic-core[mcp]'
```

### YAML configuration

Mark tools for MCP exposure in `tools.yaml`:

```yaml
# tools.yaml
tools:
  search_database:
    description: Search the database for records
    category: search
    parameters_description: "query (str): Search text"
    returns_description: "JSON with results"
    mcp:                          # NEW - MCP-specific metadata
      expose: true
      annotations:
        readOnlyHint: true

  create_record:
    description: Create a new record
    category: write
    mcp:
      expose: true
      annotations:
        readOnlyHint: false
        idempotentHint: false

  internal_tool:
    description: Internal reasoning tool
    category: reasoning
    # No mcp section → not exposed to MCP clients
```

Define which tools the MCP server exposes in `agents.yaml`:

```yaml
# agents.yaml
tool_groups:
  search_tools:
    - search_database
    - read_record

mcp_servers:
  my_server:
    description: "My knowledge base - search and manage records"
    tools:
      - group: search_tools       # reuses existing tool groups
      - list_categories
    write_tools:                  # separate list - opt-in via config
      - create_record
      - update_record
```

`description`, `tools`, `write_tools`, `resources`, and `prompts` are the whole of a server block — the server's own name comes from the mapping key — and an unrecognized key there is rejected when `get_mcp_server()` resolves it. A hyphenated `write-tools` therefore fails instead of exposing nothing. The `mcp:` block on a tool is gated the same way, so `exposed: true` fails rather than leaving the tool unexposed.

### Building the server

Implement a `MCPContextFactory` to provide runtime dependencies (database connections, auth, filters) for each tool call:

```python
from sinan_agentic_core.mcp import MCPContextFactory, build_mcp_server
from sinan_agentic_core import get_tool_registry, load_agent_catalog, load_tool_catalog

class MyContextFactory(MCPContextFactory):
    async def create_context(self):
        db = await connect_to_database()
        return MyAppContext(db_connector=db)

    async def cleanup(self, context):
        await context.db_connector.close()

# Load catalogs
agent_catalog = load_agent_catalog("agents.yaml", knowledge_dir="knowledge/")
tool_catalog = load_tool_catalog("tools.yaml")
tool_catalog.enrich_registry(get_tool_registry())

# Get MCP server config from agents.yaml
mcp_config = agent_catalog.get_mcp_server("my_server")

# Build the server
server = build_mcp_server(
    server_name="My App",
    tool_registry=get_tool_registry(),
    tool_catalog=tool_catalog,
    mcp_config=mcp_config,
    context_factory=MyContextFactory(),
    include_write_tools=False,  # read-only by default
)

# Run in stdio mode (for Claude Desktop / Claude Code)
server.run(transport="stdio")
```

### Transports

**stdio** - for local MCP clients (Claude Desktop, Claude Code). The server launches as a subprocess:

```json
{
  "mcpServers": {
    "my-app": {
      "type": "stdio",
      "command": "python",
      "args": ["-m", "my_app.mcp_server"]
    }
  }
}
```

**Streamable HTTP** - for remote clients or when sharing a process with an existing web app (e.g., FastAPI). Mount the MCP server as an ASGI app:

```python
from fastapi import FastAPI

app = FastAPI()
mcp_app = server.streamable_http_app()
app.mount("/mcp", mcp_app)
```

Clients connect via HTTP:

```json
{
  "mcpServers": {
    "my-app": {
      "type": "http",
      "url": "http://localhost:8000/mcp"
    }
  }
}
```

### How tool invocation works

The `MCPToolAdapter` bridges MCP calls to your registered `@function_tool` functions:

1. MCP client calls a tool (e.g., `search_database(query="test")`)
2. The adapter creates a context via your `MCPContextFactory`
3. It calls the function `@function_tool` decorated, reached through the SDK's
   `FunctionTool.__wrapped__`, passing the arguments straight through
4. The tool function runs with a real database connection, just like when called by an agent
5. The result is returned to the MCP client
6. The context is cleaned up (connections closed)

Each tool call gets its own context - no shared state between calls.

A direct call reaches the function *below* the SDK's function-tool pipeline, so
JSON-schema validation, tool-input guardrails, failure handling and tracing do
not run. A tool falls back to `FunctionTool.on_invoke_tool()` — a synthetic
`ToolContext` plus the arguments as JSON — whenever a direct call would change
its behavior:

- it declares tool-input guardrails, which only the SDK pipeline runs
- it exposes no function to call: a hand-built `FunctionTool`, or an agent-as-tool
- its signature declares a positional-only parameter, `*args`, or `**kwargs`,
  which the SDK maps by parameter kind

### API reference

| Function / Class | Description |
|---|---|
| `build_mcp_server(...)` | Build a FastMCP server from registry + catalogs |
| `MCPContextFactory` | ABC - implement `create_context()` and `cleanup()` |
| `MCPServerBuilder` | Low-level builder class (use `build_mcp_server()` for convenience) |
| `MCPServerConfig` | Pydantic model for resolved MCP server definition |
| `MCPToolConfig` | Per-tool MCP config (expose flag + annotations) |
| `catalog.get_mcp_tools()` | List tool names with `mcp.expose: true` |
| `catalog.get_mcp_server(name)` | Get resolved `MCPServerConfig` from `agents.yaml` |

## Knowledge Store

Inject domain knowledge into agent system prompts from separate YAML files. Knowledge teaches agents "what things are" (domain model), while tool descriptions handle routing. Inspired by the CLAUDE.md pattern - static context that shapes agent behavior.

### Setup

Create a `knowledge/` directory with one YAML file per scope:

```yaml
# knowledge/global.yaml
content: |
  The workspace is a tree of items. Pages are containers that can nest
  other pages or hold documents. Documents are imported content.

# knowledge/extraction.yaml
content: |
  Named entities are shared and deduplicated across documents.
  Insights are per-document observations with supporting evidence.
```

Reference scopes in your `agents.yaml`:

```yaml
agents:
  extraction_agent:
    model: gpt-4o
    description: Extracts structured data
    knowledge:
      - global        # -> knowledge/global.yaml
      - extraction    # -> knowledge/extraction.yaml
    tools:
      - extract_entities
```

### Loading

Pass `knowledge_dir` to `load_agent_catalog()`:

```python
from sinan_agentic_core import load_agent_catalog

catalog = load_agent_catalog("agents.yaml", knowledge_dir="knowledge/")
cfg = catalog.get("extraction_agent")
cfg.knowledge_text  # global + extraction content, concatenated with \n\n
```

Then wire it into your `AgentDefinition`:

```python
register_agent(AgentDefinition(
    name="extraction_agent",
    model=cfg.model,
    description=cfg.description,
    tools=cfg.tools,
    knowledge_text=cfg.knowledge_text,  # injected into system prompt
    instructions=build_extraction_instructions,
))
```

### How it reaches the system prompt

`InstructionBuilder.domain_knowledge()` reads `agent_def.knowledge_text` and places it between persona and context sections. The section order is:

1. **persona** - identity statement
2. **domain_knowledge** - static domain knowledge from YAML files
3. **context_section** - runtime environment info
4. **steps** - task instructions
5. **rules** - constraints
6. **output_format** - expected output structure

Override `domain_knowledge()` in your builder subclass for custom behavior.

## Structured Agent-as-Tool

When using sub-agents via the SDK's `as_tool()` pattern, you can enforce structured input and get structured error responses instead of freeform text.

### Structured input parameters

Define a dataclass for the sub-agent's input schema. The parent LLM is forced to fill in the required fields before calling the sub-agent.

```python
from dataclasses import dataclass
from sinan_agentic_core import AgentDefinition, register_agent

@dataclass
class WriterRequest:
    operation: str      # exact tool name to execute
    target_id: str      # primary target identifier
    payload: str = ""   # additional params as JSON

register_agent(AgentDefinition(
    name="writer_agent",
    description="Execute write operations on the database.",
    instructions="You are a write-operation executor.",
    tools=["create_record", "update_record", "delete_record"],
    as_tool_parameters=WriterRequest,  # enforced input schema
))
```

When the parent agent calls `writer_agent`, the LLM must provide `operation`, `target_id`, and optionally `payload` - no more ambiguous freeform text.

### Structured error handling

Sub-agent failures automatically return structured JSON instead of generic error strings:

```json
{
    "status": "error",
    "error_type": "ValueError",
    "message": "target_id is required",
    "retry_hint": "A required parameter is missing. Check your context for available IDs and provide all required fields."
}
```

This is handled by `structured_tool_error()`, which is automatically wired into all agent-as-tool calls via `BaseAgentRunner`. The hint is picked from the failure's *type* first — `classify_run_error()` maps the exception onto a `RunErrorKind` — and only falls back to the message for the `ValueError`s this framework's own tools raise:

| Failure | Retry hint |
|---------|------------|
| `MaxTurnsExceeded` | Simplify the request or break it into smaller steps |
| Provider `context_length_exceeded` | Narrow the request, or split it across several calls |
| `ModelRefusalError` | Do not re-send the same request — restate it or handle it another way |
| `ModelBehaviorError` | The response missed its output schema; retry once with a simpler request |
| `InputGuardrailTripwireTriggered` | A guardrail rejected the request; do not re-send it |
| `OutputGuardrailTripwireTriggered` | A guardrail blocked the answer; report it rather than retrying |
| "not found" in message | Verify the ID exists in your context |
| "required" in message | Check context for available IDs and provide all required fields |
| Other | Review the error message and retry with corrected input |

Keying off the class rather than the wording means an upstream release that rewords `"Max turns (10) exceeded"` does not silently downgrade the hint, and an unrelated error that merely quotes those words is not mistaken for the failure it names.

### How it works

`BaseAgentRunner._build_tools()` automatically:
1. Passes `as_tool_parameters` to `agent.as_tool(parameters=...)` when defined
2. Passes `structured_tool_error` as `failure_error_function` for all agent-as-tool calls
3. If `as_tool_turn_budget` is set, wires up the sub-agent's own hooks and steering (see [Turn budget for sub-agents](#turn-budget-for-sub-agents-agent-as-tool))

No manual wiring needed - just set `as_tool_parameters` on your `AgentDefinition`.

## Extending sinan-agentic — Custom Capabilities

`Capability` is the extension point for cross-cutting agent behavior. Write a subclass, attach it to an `AgentDefinition`, and the runtime calls its lifecycle hooks (and appends its instruction fragments to the model input) for every run. `TurnBudget`, `ToolErrorRecovery`, and `ToolTracer` are themselves `Capability` subclasses - see them for reference.

```python
from agents import RunContextWrapper, Tool
from sinan_agentic_core import Capability, AgentDefinition, register_agent


class ToolCallLogger(Capability):
    """Print every tool call - useful as a debugging aid."""

    def on_tool_start(self, ctx: RunContextWrapper, tool: Tool, args: str) -> None:
        print(f"[tool] -> {tool.name}({args})")

    def on_tool_end(self, ctx: RunContextWrapper, tool: Tool, result: str) -> None:
        print(f"[tool] <- {tool.name}: {result[:120]}")


register_agent(AgentDefinition(
    name="time_assistant",
    description="Tells the time",
    instructions="You are a helpful assistant.",
    tools=["get_current_time"],
    capabilities=[ToolCallLogger()],
))
```

The runner clones each declared capability per run (so concurrent runs are isolated), calls `reset()` on the clones, then composes them into a single `RunHooks` adapter. Override only the hook methods you care about - everything else is a no-op by default.

For the full lifecycle reference, the protocol surface, and migration notes for existing direct-Python and YAML users, see [`documentation/project/capabilities.md`](documentation/project/capabilities.md). For a runnable end-to-end example, see [`examples/custom_capability.py`](examples/custom_capability.py).

## Tool Error Recovery

When a tool returns an error, agents often retry with identical parameters, wasting turns in a loop. `ToolErrorRecovery` solves this by tracking tool errors and steering the next model call with progressive recovery guidance - the same capability-steering pattern used by `TurnBudget`.

### How it works

1. **`on_tool_end` hook** tracks tool results. When a result contains `{"error": "..."}`, the tool name, arguments, and error are recorded.
2. **Capability steering** appends a `## Tool Error Recovery` section to the model input before each LLM call, telling the agent what failed and how to recover.
3. **Progressive escalation** - guidance gets more directive with each repeated failure:
   - **1st failure**: Show the error and recovery hint
   - **2nd failure (same args)**: Warn that the agent is repeating the same failing call
   - **3rd failure (same args)**: Tell the agent to stop retrying and move on

### Registering recovery hints

Add a `recovery_hint` to your tool registration. This static hint is shown to the agent whenever the tool errors - no need to handle each error case individually:

```python
@register_tool(
    name="search_database",
    description="Search the database by query",
    category="search",
    parameters_description="query (str): Search query. scope (str): 'local' or 'global'.",
    returns_description="JSON with results",
    recovery_hint="Requires a non-empty query. Use scope='local' for fast search, 'global' for external APIs.",
)
@function_tool
async def search_database(ctx, query: str, scope: str = "local") -> str:
    ...
```

For MCP tools (where you don't control the code), pass hints via config:

```python
recovery = ToolErrorRecovery(
    tool_registry=registry,
    mcp_hints={
        "mcp_arxiv_search": "Query must be at least 2 characters.",
        "mcp_slack_post": "Channel ID is required, not channel name.",
    },
)
```

### Usage with BaseAgentRunner

Pass `error_recovery` to `execute()`:

```python
from sinan_agentic_core import BaseAgentRunner, ToolErrorRecovery

runner = BaseAgentRunner()
recovery = ToolErrorRecovery(tool_registry=runner.tool_registry)

output = await runner.execute(
    agent_name="my_agent",
    context=context,
    session=session,
    input_text="Find recent papers on transformers",
    error_recovery=recovery,
)

# After execution, inspect error state:
if recovery.has_errors:
    print(recovery.get_error_summary())
```

### Combining with Turn Budget

Both features compose automatically. When both are active, the agent sees both sections in its instructions:

```python
from sinan_agentic_core import BaseAgentRunner, TurnBudget, ToolErrorRecovery

runner = BaseAgentRunner()
budget = TurnBudget(default_turns=15)
recovery = ToolErrorRecovery(tool_registry=runner.tool_registry)

output = await runner.execute(
    agent_name="research_agent",
    context=context,
    session=session,
    input_text="Analyze recent publications",
    turn_budget=budget,
    error_recovery=recovery,
)
```

The runner composes both hook sets into a single `_CompositeHooks` and joins both steering sections into the one trailing input item.

### What the agent sees

After a tool error, the input the agent is sent ends with:

```text
## Tool Error Recovery
- search_database returned error: "No results found" (called with: query=transformrs, scope=local)
  Recovery hint: Requires a non-empty query. Use scope='local' for fast search, 'global' for external APIs.

General rule: Never retry a tool call with identical parameters. If a tool fails, read the error and choose a different approach.
```

After repeated identical failures:

```text
## Tool Error Recovery
- STOP: search_database has failed 3 times with the same arguments. Do NOT call this tool again.
  Error was: "Connection timeout"
  Move on to your next task, try a completely different approach, or return your partial results.
```

### YAML configuration

Recovery is on by default for every YAML agent. The `error_recovery:` shorthand turns it off, or on with the defaults; the `capabilities:` list turns it on and tunes it:

```yaml
agents:
  simple_agent:
    model: gpt-4o-mini
    description: Quick tasks
    error_recovery: false        # off — failures surface untouched

  mcp_agent:
    model: gpt-4o
    description: Drives external MCP servers
    error_recovery: false        # the shorthand is on by default, and both
                                 # forms apply — leave it on and the agent
                                 # gets two recovery capabilities
    capabilities:
      - name: error_recovery
        config:
          max_identical_before_stop: 5
          mcp_hints:
            mcp_arxiv_search: Query must be at least 2 characters.
```

Those two are the whole of that config block, and an unrecognized key there is rejected when the capability is built. `tool_registry` is not one of them: it is a live object rather than something YAML can express, and it falls back to the process-wide registry.

### Configuration reference

| Parameter | Default | Description |
|-----------|---------|-------------|
| `tool_registry` | None | ToolRegistry for looking up `recovery_hint` |
| `mcp_hints` | `{}` | Dict mapping MCP tool names to hint strings |
| `max_identical_before_stop` | 3 | Stop threshold for identical-argument retries |

## Structured Output Recovery

An agent with an `output_dataclass` fails its whole run when the model's final message does not parse — including when the payload itself is fine but the model wrapped it in a ```` ```json ```` fence or a "Here is the result:" preamble. The SDK hands the raw message straight to Pydantic, so those runs raise `ModelBehaviorError` even though the answer is right there.

Every run started through `BaseAgentRunner` or the chat functions installs a recovery handler that re-reads the message, finds the embedded payload, and validates it against the agent's own output schema. It never fabricates a value: anything that does not satisfy the schema is not recovered, and the run raises as before.

Session history reads messages the same way. An agent with a non-dict `output_dataclass` answers under a wrapper key, and `AgentSession` stores the answer rather than the envelope — including when the model fenced its payload or wrote a preamble around it, which previously left the fence markers and the wrapper in the conversation every later turn replays.

Recovery is on by default and covers `execute()` in all three modes plus the overflow-fallback LLM call. `invalid_output_recovery` is an agent-definition flag, so it turns recovery off everywhere that definition is resolved — the runner's branches, the chat functions, and a `Runner.run()` you drive yourself. Turn it off per agent when a malformed response must fail loudly:

```yaml
# agents.yaml
agents:
  extractor:
    model: gpt-4o-mini
    description: Extracts structured records
    invalid_output_recovery: false
```

```python
register_agent(AgentDefinition(
    name="extractor",
    description=cfg.description,
    instructions=build_extractor_instructions,
    output_dataclass=ExtractionResult,
    invalid_output_recovery=cfg.invalid_output_recovery,
))
```

Running `Runner.run()` directly instead of through `BaseAgentRunner`? `build_error_handlers()` reads the decision the way `build_run_config()` does — the output type off the agent, and the `invalid_output_recovery` flag off the definition registered under the agent's name, since the flag has no slot on `Agent`. An agent that answers in plain text has no schema to fail, and an agent whose definition turns recovery off asked to fail loudly; both get `None` and `Runner.run()` keeps its defaults:

```python
from agents import Runner
from sinan_agentic_core import build_error_handlers

result = await Runner.run(
    agent,
    "Extract the records",
    error_handlers=build_error_handlers(agent),
)
```

The chat functions do exactly this on the agent they resolve. Pass `recover_invalid_final_output` yourself when you assemble the handler mapping by hand.

Sub-agents built with `as_tool()` are the one gap — the SDK's `as_tool()` has no `error_handlers` parameter, so a nested agent's invalid structured output still raises.

## Model Retry Policies

A rate limit, a 503, or a dropped connection fails the whole run — even though the next attempt would have succeeded. The SDK can retry the model call itself, but only when `ModelSettings.retry` carries both an attempt budget *and* a policy callback; with either missing it never retries.

Declare `model_retry` on an agent and the framework builds that object for you. It reaches every execution path — `execute()` in all three modes, `run_agent()`, handoffs, and `as_tool()` sub-agents — because retry rides on the agent's model settings rather than on a per-run argument.

Retry is off unless declared: a retried call costs latency and a second billed request.

```yaml
# agents.yaml
agents:
  researcher:
    model: gpt-4o-mini
    description: Reads papers and summarizes findings
    model_retry:
      max_retries: 3               # attempts after the initial request
      retry_on: [provider_suggested, network_error]
      backoff:                     # optional — SDK defaults fill any unset field
        initial_delay: 0.5
        max_delay: 8.0
```

Every field is optional, so `model_retry: {}` opts in with the defaults — two attempts after the initial request, on the triggers listed below. Only a missing key means the agent opts out.

`max_retries`, `retry_on`, and `backoff` are the whole of the `model_retry:` block, and an unrecognized key — there or inside `backoff:` — is rejected at load time. A delay field written beside `backoff:` instead of under it therefore fails, rather than leaving the schedule at SDK defaults while reading as declared.

```python
from sinan_agentic_core import AgentDefinition, ModelRetryConfig, register_agent

register_agent(AgentDefinition(
    name="researcher",
    description=cfg.description,
    instructions=build_researcher_instructions,
    model_retry=cfg.model_retry,          # from agents.yaml
    # ...or inline: ModelRetryConfig(max_retries=3)
))
```

`retry_on` lists the error classes worth another attempt. Defaults to `[provider_suggested, network_error]`.

| Trigger | Retries when |
|---------|--------------|
| `provider_suggested` | The model adapter advises it. For OpenAI: 408, 409, 429, every 5xx, connection errors and timeouts — honoring an `x-should-retry` header or a `Retry-After` delay when one is sent. |
| `network_error` | The call failed with a connection error or a timeout, whatever the provider advises. |
| `retry_after` | The error carries an explicit `Retry-After` delay, and waits exactly that long. |

Triggers combine: the call is retried when any listed trigger matches. A delay supplied by the error (`Retry-After`) always wins over the `backoff` schedule.

Both agent-building paths attach the declared policy the same way: `BaseAgentRunner.create_agent()` and `create_agent_from_registry()`. An agent resolved by name in `chat()`, `chat_with_hooks()`, or `chat_streamed()` therefore retries too. Settings supplied by the caller still win field by field, so an explicit `retry` on a `model_settings` override replaces the declared one.

The overflow-fallback path is the one partial case. Its rescue call goes straight to the OpenAI client instead of through the runner, so it honors `max_retries` but not `retry_on` or `backoff`.

## Model Call Timeout

Nothing bounds a model call by default. A provider that accepts the request and then stalls holds the run open until the caller gives up, and a declared `model_retry` makes that worse rather than better — an attempt budget multiplies an unbounded wait instead of capping it.

Declare `model_timeout` on an agent and each model-call attempt gets a deadline in seconds. Like `model_retry`, it rides on the agent's model settings rather than on a per-run argument, so it reaches every execution path — `execute()` in all three modes, `run_agent()`, handoffs, and `as_tool()` sub-agents.

```yaml
# agents.yaml
agents:
  researcher:
    model: gpt-4o-mini
    description: Reads papers and summarizes findings
    model_timeout: 30          # seconds, per attempt
    model_retry:
      max_retries: 3           # ...so three attempts are bounded at 30s each
```

`model_timeout` is its own key, not part of the `model_retry:` block. Bounding how long one attempt may hang and deciding whether to pay for another one are separate choices, and either is useful without the other. The value is seconds and must be greater than zero; a non-positive or infinite one is rejected at load time.

```python
from sinan_agentic_core import AgentDefinition, register_agent

register_agent(AgentDefinition(
    name="researcher",
    description=cfg.description,
    instructions=build_researcher_instructions,
    model_timeout=cfg.model_timeout,      # from agents.yaml
    # ...or inline: model_timeout=30.0
))
```

The bound covers the complete attempt, transport waits included, and is enforced through normal asyncio cancellation. It does **not** bound the whole run, tool calls, or the delay a `backoff` schedule waits between attempts — `turn_budget` and `max_turns` are what cap the loop itself.

When the deadline passes the run fails with the SDK's `ModelTimeoutError`, which `classify_run_error()` reports as `RunErrorKind.MODEL_TIMEOUT` and `chat()` returns as `error_kind: "model_timeout"`. That kind exists so a caller can tell their own bound firing apart from an unrecoverable failure — the answer to one is usually a larger number, not a bug hunt.

Timing out is not one of the failures `execute(fallback_on_overflow=True)` rescues: the rescue is another model call, bounded by the same number of seconds, so a provider slow enough to trip the bound trips it again. The rescue call is bounded — it goes straight to the OpenAI client, which gets the declared timeout as its own per-request limit.

Both agent-building paths attach the bound the same way: `BaseAgentRunner.create_agent()` and `create_agent_from_registry()`. Settings supplied by the caller still win field by field, so an explicit `timeout` on a `model_settings` override replaces the declared one.

## Tool Output Trimming

A long **multi-turn** conversation accumulates bulky tool outputs — search results, file dumps, error payloads — that keep costing tokens long after they stop being relevant. Left alone they grow the input until the run dies on `context_length_exceeded`, and the only recovery left is `execute(fallback_on_overflow=True)`, which re-collects the tool outputs and pays for a second, condensed model call.

Declare `tool_output_trim` on an agent and the SDK replaces oversized tool outputs from older turns with a short preview *before* each model call. Recent turns stay at full fidelity, so the agent keeps the context it is actually working with while the tail stops growing. It reaches every execution path that runs through the SDK — `execute()` in all three modes, `run_agent()`, `as_tool()` sub-agents, and the chat functions.

Trimming is off unless declared: it removes content the model would otherwise have seen.

**It only helps across turns.** The window is counted in *user messages*, so nothing is trimmed until the conversation holds more than `recent_turns` of them. A single question answered by twenty large tool calls is left whole at every setting — the outputs all sit in the current turn, and the current turn is never touched. Trimming pays off for a session that keeps asking follow-ups, not for a one-shot run that overflows inside its own turn; that shape still needs `fallback_on_overflow`.

```yaml
# agents.yaml
agents:
  researcher:
    model: gpt-4o-mini
    description: Reads papers and summarizes findings
    tool_output_trim:
      recent_turns: 3              # last 3 user turns are never trimmed
      max_output_chars: 4000       # outputs above this are candidates
      preview_chars: 500           # how much of the original survives
      trimmable_tools: [web_search]  # optional — omit to cover every tool
```

Every field is optional and the SDK fills an unset one with its own default, so `tool_output_trim: {}` opts in with defaults throughout.

Those four are the whole of the `tool_output_trim:` block, and an unrecognized key there is rejected at load time — a misspelled `max_output_char` fails instead of trimming on the SDK's default cap.

```python
from sinan_agentic_core import AgentDefinition, ToolOutputTrimConfig, register_agent

register_agent(AgentDefinition(
    name="researcher",
    description=cfg.description,
    instructions=build_researcher_instructions,
    tool_output_trim=cfg.tool_output_trim,   # from agents.yaml
    # ...or inline: ToolOutputTrimConfig(max_output_chars=4000)
))
```

The SDK installs the filter through `RunConfig`, so unlike a retry policy it cannot ride on the agent — building one is not enough to get it. `build_run_config()` closes that gap: it looks the declaration up by the agent's name, so an agent resolved in `chat()`, `chat_with_hooks()`, or `chat_streamed()` is trimmed on the same terms as one the runner executes. Driving `Runner` yourself works the same way:

```python
from agents import Runner
from sinan_agentic_core import build_run_config, create_agent_from_registry

agent = create_agent_from_registry("researcher")
result = await Runner.run(agent, "Summarize the new papers", run_config=build_run_config(agent))
```

An agent assembled by hand under a name the registry does not know declares nothing and keeps the SDK defaults.

Pick `max_output_chars` above the size of an output the agent still reasons over several turns later, and `preview_chars` large enough that the trimmed entry still says what the call returned. The two are independent: an output is replaced only when the preview is genuinely shorter than the original, so `preview_chars: 500` with `max_output_chars: 500` is a valid "keep the first 500 characters of anything bigger" policy — it spares outputs just over the cap and still cuts a 100,000-character one by 99%.

Trimming and the overflow fallback are complements, not alternatives: trimming makes overflow rarer, the fallback still rescues the run when it happens anyway. The fallback's own rescue call bypasses the SDK, so a declared trim policy does not shape that prompt — the prompt builder caps each output instead.

The fallback fires on exactly two failures — the SDK's `MaxTurnsExceeded` and a provider 400 carrying `context_length_exceeded`. Both mean the loop ran out of room. Everything else propagates, a `ModelRefusalError` included: the rescue call goes straight to the model and re-asking a refusal through a path that bypasses the run would route around the answer rather than recover from a limit.

## Tool Tracer (Non-Streaming Observability)

The streaming path emits per-tool events through `on_event`, but non-streaming runs (`streaming=False`, the default) have no equivalent channel. `ToolTracer` is a built-in `Capability` that prints one line per tool call and agent boundary, so non-streaming consumers can see what the agent is doing without leaving the framework or reading `session.history` after the fact.

### Usage

Attach it like any other capability — declarative on the agent, or runtime through `execute()`:

```python
from sinan_agentic_core import AgentDefinition, ToolTracer, register_agent

register_agent(AgentDefinition(
    name="research_agent",
    description="Reads papers and summarizes findings",
    instructions="...",
    tools=["search_papers", "fetch_url"],
    capabilities=[ToolTracer()],
))
```

Or pass a custom sink (any `Callable[[str], None]`) — for example, a logger:

```python
import logging

logger = logging.getLogger("agent.trace")
tracer = ToolTracer(sink=logger.info, truncate_args=120, truncate_result=300)
```

Each run emits lines such as:

```text
[agent start] research_agent
[tool start] search_papers args={"query": "transformer scaling laws"}
[tool end] search_papers result=[{"id": "2401.12345", ...
[agent end] research_agent
```

### YAML

The `tool_tracer` factory is pre-registered in `CapabilityRegistry`, so YAML agents can opt in with no Python:

```yaml
agents:
  research_agent:
    model: gpt-4o
    description: Reads papers and summarizes findings
    capabilities:
      - name: tool_tracer
        config:
          truncate_args: 120
          truncate_result: 300
          include_timestamps: true
```

`sink` is not configurable from YAML — it falls back to `print`. To pipe trace lines into a logger, attach the tracer in code instead.

### Configuration reference

| Parameter | Default | Description |
|-----------|---------|-------------|
| `sink` | `print` | One-arg callable receiving each line |
| `truncate_args` | `200` | Maximum length of the rendered `args` string (`0` disables) |
| `truncate_result` | `500` | Maximum length of the rendered `result` string (`0` disables) |
| `include_timestamps` | `False` | Prepend `HH:MM:SS.mmm` wall-clock prefix to every line |

`ToolTracer` is observation-only and never mutates state used by the agent, so it composes cleanly with `TurnBudget` and `ToolErrorRecovery` on the same agent.

## Turn Budget (Soft Turn Management)

The OpenAI Agents SDK enforces `max_turns` as a hard cutoff -- when hit, everything stops with no graceful handling. `TurnBudget` adds a soft budget layer where the agent self-manages its turns:

- **Default budget** -- the agent perceives a soft turn limit (e.g., 10 turns)
- **Hard ceiling** -- the SDK's `max_turns` is set to a higher absolute maximum (safety net)
- **Warnings** -- the agent gets notified when running low on turns
- **Self-extension** -- the agent can call `request_extension()` to approve more turns for itself

### Concept

```text
TurnBudget(default_turns=10)
  |
  |-- Agent perceives 10 turns (soft limit)
  |-- SDK gets max_turns=25 (hard ceiling, invisible to agent)
  |-- At turn 8: "2 turns remaining, wrap up or request_extension()"
  |-- Agent calls request_extension("processing remaining docs") -> budget extends to 15
  |-- If truly exhausted: error handler provides graceful fallback
```

### YAML configuration

Define turn budgets per agent in `agents.yaml`. The `max_turns` field becomes the hard ceiling (`absolute_max`), and `turn_budget` controls the soft budget:

```yaml
agents:
  research_agent:
    model: gpt-4o
    max_turns: 25            # hard ceiling (SDK safety net)
    description: Deep research
    tools: [think, search, read]
    turn_budget:
      default_turns: 10      # soft limit the agent perceives
      reminder_at: 2         # warn when 2 turns remain
      max_extensions: 3      # agent can self-extend up to 3 times
      extension_size: 5      # turns added per extension

  simple_agent:
    model: gpt-4o-mini
    max_turns: 10
    description: Quick tasks
    tools: [think]
    # No turn_budget -- uses plain max_turns cutoff
```

Every field is optional, so `turn_budget: {}` opts in with the defaults above. Only a missing key means the agent opts out.

Then use `build_turn_budget()` to create a `TurnBudget` from the catalog entry:

```python
from sinan_agentic_core import load_agent_catalog, BaseAgentRunner

catalog = load_agent_catalog("agents.yaml")
cfg = catalog.get("research_agent")
budget = cfg.build_turn_budget()  # TurnBudget or None if not configured

runner = BaseAgentRunner()
output = await runner.execute(
    agent_name="research_agent",
    context=context,
    session=session,
    input_text="Analyze these 10 papers",
    max_turns=cfg.max_turns,
    turn_budget=budget,
)
```

### Programmatic usage

You can also create a `TurnBudget` directly:

```python
from sinan_agentic_core import BaseAgentRunner, TurnBudget

budget = TurnBudget(
    default_turns=10,
    reminder_at=2,
    max_extensions=3,
    extension_size=5,
    absolute_max=25,
)

output = await runner.execute(
    agent_name="research_agent",
    context=context,
    session=session,
    input_text="Analyze these 10 papers",
    turn_budget=budget,
)

# After execution, inspect budget state:
print(budget.turns_used)        # 14
print(budget.extensions_used)   # 1
print(budget.extension_reasons) # ["Need to process 5 remaining papers"]
```

### How the budget reaches the model

When a `TurnBudget` is provided, the runner:

1. Sets the SDK's `max_turns` to `budget.absolute_max` (hard safety ceiling)
2. Carries `budget` on the run context via `set_turn_budget()` (read back with `get_turn_budget()`)
3. Adds the `request_extension` tool to the agent
4. Appends budget status to the model input before each call, through the run config's `call_model_input_filter`
5. Counts turns through the capability's `on_llm_start` hook

If using `InstructionBuilder`, the `turn_budget_section()` method is included in the default section order and automatically reads the budget from context.

### Turn budget for sub-agents (agent-as-tool)

Sub-agents running via `.as_tool()` can have their own independent turn budget. Set `as_tool_turn_budget` on the `AgentDefinition` instead of `as_tool_max_turns`:

```yaml
agents:
  writer_agent:
    model: gpt-4o-mini
    max_turns: 10
    turn_budget:
      default_turns: 5
      reminder_at: 1
      max_extensions: 1
      extension_size: 3
```

```python
cfg = catalog.get("writer_agent")

register_agent(AgentDefinition(
    name="writer_agent",
    description="Execute write operations on the database.",
    instructions="You are a write-operation executor.",
    tools=["create_record", "update_record", "delete_record"],
    as_tool_turn_budget=cfg.build_turn_budget(),  # budget instead of as_tool_max_turns
))
```

When `as_tool_turn_budget` is set, `_build_tools()` automatically:
1. Resets the budget for each parent agent creation
2. Steers the sub-agent's own model calls with its budget status, through the run config `.as_tool(run_config=...)` receives
3. Adds the `request_extension` tool to the sub-agent
4. Passes the sub-agent's hooks to `.as_tool(hooks=...)` for turn tracking
5. Sets `max_turns` to `budget.absolute_max` (overrides `as_tool_max_turns`)

If `as_tool_turn_budget` is not set, the runner falls back to `as_tool_max_turns` (plain hard cap).

### Configuration reference

| Parameter | Default | Description |
|-----------|---------|-------------|
| `default_turns` | 10 | Soft budget the agent perceives |
| `reminder_at` | 2 | Warn when this many turns remain |
| `max_extensions` | 3 | Max self-approved extensions |
| `extension_size` | 5 | Turns added per extension |

Those four are the whole of the `turn_budget:` block, and an unrecognized key
there is rejected at load time. The hard ceiling — `TurnBudget.absolute_max`, 25
when unset — is not one of them: in YAML it is the agent's own `max_turns`, and
in code it is a constructor argument.

## Session Persistence

```python
from sinan_agentic_core import AgentSession, SQLiteSessionStore

# In-memory (default)
session = AgentSession(session_id="user-123")
await session.add_items([{"role": "user", "content": "Hello!"}])
history = await session.get_items()

# SQLite (persistent)
store = SQLiteSessionStore("data/conversations.db")
store.add_message("session-123", "user", "What's the weather?")
store.add_message("session-123", "assistant", "It's sunny.")
history = store.get_conversation_history("session-123")  # [{"role": ..., "content": ...}]
store.archive_session("session-123")   # Archive (keeps data)
store.clear_session("session-123")     # Delete permanently
```

## Output Models

```python
from sinan_agentic_core import ToolOutput, ChatResponse

output = ToolOutput(success=True, data={"temp": 72}, metadata={"source": "api"})
response = ChatResponse(success=True, response="It's sunny.", session_id="u-123", tools_called=["get_weather"])

output.to_dict()    # {"success": True, "data": {...}, "metadata": {...}}
response.to_dict()  # {"success": True, "response": "...", ...}
```

## Project Structure

```text
sinan_agentic_core/
├── __init__.py              # Main exports
├── orchestrator.py          # Multi-agent orchestration
├── core/
│   ├── base_runner.py            # BaseAgentRunner
│   ├── capabilities/             # Capability protocol (pluggable agent behaviors)
│   │   ├── __init__.py
│   │   ├── base.py               # Capability base class + lifecycle hooks
│   │   └── steering.py           # Delivers instruction fragments at the tail of the model input
│   ├── errors.py                 # structured_tool_error for agent-as-tool failures
│   ├── output_recovery.py        # invalid_final_output handler (salvages structured output)
│   ├── run_config.py             # build_run_config() — run-level SDK settings for a built agent
│   ├── tool_error_recovery.py    # ToolErrorRecovery capability
│   ├── tool_tracer.py            # ToolTracer capability (non-streaming tool-call tracing)
│   ├── turn_budget.py            # TurnBudget capability
│   └── turn_budget_tool.py       # request_extension tool (agent self-approval)
├── instructions/
│   └── builder.py           # InstructionBuilder base class
├── mcp/                        # MCP server support (optional, requires sinan-agentic-core[mcp])
│   ├── __init__.py             # Public API: build_mcp_server, MCPContextFactory
│   ├── context_protocol.py     # MCPContextFactory ABC
│   ├── server_builder.py       # MCPServerBuilder + build_mcp_server()
│   ├── tool_adapter.py         # Wraps registered tools for MCP invocation
│   └── yaml_schema.py          # Pydantic models for MCP YAML config
├── registry/
│   ├── agent_catalog.py        # YAML-driven agent catalog + MCP server definitions
│   ├── agent_registry.py       # AgentDefinition + registry
│   ├── agent_factory.py        # create_agent_from_registry()
│   ├── capability_registry.py  # @register_capability + named capability factories
│   ├── tool_catalog.py         # YAML-driven tool catalog + MCP exposure flags
│   ├── tool_registry.py        # ToolDefinition + registry
│   └── guardrail_registry.py
├── services/
│   ├── chat.py              # chat(), chat_with_hooks(), chat_streamed()
│   ├── hooks.py             # StreamingRunHooks (tool call tracking)
│   └── events.py            # Event dataclasses + StreamingHelper
├── session/
│   ├── agent_session.py     # In-memory session
│   └── sqlite_store.py      # SQLite persistence
└── models/
    ├── context.py          # AgentContext
    └── outputs/            # ToolOutput, ChatResponse
```

## Future Work

### Conversation-Driven Learning (Lessons)

Agents today are stateless across sessions. When an agent makes a mistake, the only fix is manually adding rules to instructions, which causes instruction bloat and degrades attention over time.

The goal: a post-conversation analysis pipeline that automatically extracts lessons from completed conversations and retrieves relevant ones at the start of future sessions.

**Concept:**
- An analysis agent reads completed conversation transcripts
- It detects mistakes, inefficiencies, and user corrections (explicit: "no, I meant X"; implicit: wasted tool calls, repeated searches)
- It extracts structured lessons (trigger context, mistake, correct behavior, confidence)
- Lessons are stored in a `LessonStore` (SQLite default, overridable)
- At conversation start, `InstructionBuilder` retrieves relevant lessons by context similarity and injects them as dynamic context

**Key difference from rules:** rules grow linearly in the prompt; lessons scale in a database and only the 2-3 relevant ones are injected per conversation.

**Open problems:**
- Signal detection: how to reliably identify that a conversation went badly (especially when mistakes are subtle)
- Generalization: turning one specific mistake into a reusable lesson without overfitting
- Retrieval: selecting the right lessons from hundreds, with minimal context at conversation start
- Quality: preventing bad lessons from degrading future performance (confidence scoring, decay, contradiction detection)

**Research references:** ExpeL (cross-task insight extraction), Memento (case-based reasoning), A-Mem (self-organizing memory), Reflexion (verbal self-reflection), DSPy (automatic prompt optimization). See `documentation/04-brainstorms/10-agent-learning-from-conversations.md` in Digital Brain for the full analysis.

### Dynamic System Prompts

The OpenAI SDK re-sends the full system prompt every turn. For long instructions, this wastes tokens and may cause re-deliberation (agent re-reads its full playbook and reconsiders strategy mid-execution).

**Possible approaches:**
- Turn-aware instructions: full prompt on turn 1, condensed on turn 2+ (knowledge already in message history)
- LLM-rewritten instructions: dynamically condense instructions based on conversation state
- ITR-style retrieval: per-turn RAG over instruction fragments and tool subsets (95% token reduction in research, arxiv 2602.17046)

**Status:** Parked. Current instructions work and there's no evidence of real problems. Worth revisiting when instruction size or conversation length becomes a measurable bottleneck.

## License

MIT License
