Metadata-Version: 2.4
Name: toolproxy
Version: 0.6.0
Summary: Universal tool-calling wrapper for any LLM — native & emulated function calling, argument auto-repair, usage budgets, loop detection, caching, retry, persistent SQLite memory, structured output, and FastAPI integration
Project-URL: Homepage, https://github.com/Dheena731/toolproxy
Project-URL: Repository, https://github.com/Dheena731/toolproxy
Project-URL: Bug Tracker, https://github.com/Dheena731/toolproxy/issues
Project-URL: Changelog, https://github.com/Dheena731/toolproxy/blob/main/README.md#changelog
Project-URL: Documentation, https://dheena731.github.io/toolproxy/
Author: toolproxy contributors
License: MIT
License-File: LICENSE
Keywords: agent,agents,ai,anthropic,async,caching,fastapi,function-calling,gemini,groq,llm,mcp,memory,model-context-protocol,ollama,openai,openrouter,pydantic,retry,streaming,structured-output,tool-calling
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Internet :: WWW/HTTP :: HTTP Servers
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: httpx>=0.25
Requires-Dist: openai>=1.0
Requires-Dist: pydantic>=2.0
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.25; extra == 'anthropic'
Provides-Extra: cohere
Provides-Extra: dev
Requires-Dist: build>=1.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.21; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest-mock>=3.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: twine>=5.0; extra == 'dev'
Provides-Extra: fastapi
Requires-Dist: fastapi>=0.110; extra == 'fastapi'
Requires-Dist: httpx>=0.25; extra == 'fastapi'
Provides-Extra: full
Requires-Dist: anthropic>=0.25; extra == 'full'
Requires-Dist: fastapi>=0.110; extra == 'full'
Requires-Dist: langchain-core>=0.1; extra == 'full'
Requires-Dist: opentelemetry-api>=1.20; extra == 'full'
Requires-Dist: opentelemetry-sdk>=1.20; extra == 'full'
Provides-Extra: gemini
Provides-Extra: groq
Provides-Extra: langchain
Requires-Dist: langchain-core>=0.1; extra == 'langchain'
Provides-Extra: otel
Requires-Dist: opentelemetry-api>=1.20; extra == 'otel'
Requires-Dist: opentelemetry-sdk>=1.20; extra == 'otel'
Description-Content-Type: text/markdown

# toolproxy

[![PyPI version](https://badge.fury.io/py/toolproxy.svg)](https://pypi.org/project/toolproxy/)
[![Python Versions](https://img.shields.io/pypi/pyversions/toolproxy.svg)](https://pypi.org/project/toolproxy/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Tests](https://img.shields.io/badge/tests-396%20passing-brightgreen)](https://github.com/Dheena731/toolproxy)

**Universal Tool-Calling Wrapper for Non-Tool-Native LLMs**

A provider-agnostic Python library that adds reliable tool/function calling to *any* LLM — even models with no native tool-calling support. One API, every model, sync and async.

```python
from toolproxy import UniversalAgent, tool

@tool
def get_weather(city: str) -> str:
    """Get the current weather for a city."""
    return f"Sunny, 25°C in {city}"

agent = UniversalAgent(model="openrouter/mistralai/mistral-7b-instruct", tools=[get_weather])
result = agent.run("What's the weather in Chennai?")
print(result.content)  # → "The weather in Chennai is Sunny, 25°C."
```

---

## Why toolproxy?

| Problem | toolproxy's Solution |
|---|---|
| My model doesn't support function calling | **Emulated mode** — injects a structured JSON protocol into the system prompt |
| I want native tool calls when available | **Auto-detection** — switches transparently based on model capability |
| Switching providers breaks my code | **Single unified API** — same `agent.run()` for every model |
| Tool calls hit the same API repeatedly | **Result caching** — skip redundant calls, configurable TTL |
| Network blips crash my agent | **Retry with backoff** — automatic recovery from transient failures |
| I lose context between sessions | **Persistent memory** — JSON file-backed history that survives restarts |
| I want to serve my agent as an HTTP API | **FastAPI integration** — one line to expose run/stream/reset endpoints |
| The model sends malformed/misspelled arguments | **Argument auto-repair** — fuzzy keys, JSON unwrapping, type coercion before validation *(v0.6)* |
| A runaway agent quietly racks up a huge bill | **Usage budgets** — cap tool calls and tokens per run; `result.usage` accounting *(v0.6)* |
| The model gets stuck calling the same tool forever | **Loop detection** — abort on repeated identical calls *(v0.6)* |

---

## Installation

```bash
# Core (no optional deps)
pip install toolproxy

# With FastAPI integration
pip install "toolproxy[fastapi]"

# With Anthropic adapter
pip install "toolproxy[anthropic]"

# Everything
pip install "toolproxy[full]"
```

**Requires Python 3.10+**

---

## Table of Contents

- [Quick Start](#quick-start)
- [How It Works](#how-it-works)
- [Supported Models & Adapters](#supported-models--adapters)
- [Core API](#core-api)
  - [The `@tool` Decorator](#the-tool-decorator)
  - [UniversalAgent](#universalagent)
  - [run() and arun()](#run-and-arun)
  - [Streaming](#streaming)
- [v0.4 Features](#v04-features)
  - [Tool Result Caching](#tool-result-caching)
  - [Retry with Backoff](#retry-with-backoff)
  - [Persistent File Memory](#persistent-file-memory)
  - [Agent Extensions](#agent-extensions)
  - [FastAPI Integration](#fastapi-integration)
- [v0.5 Features](#v05-features)
  - [Parallel Emulated Tools](#parallel-emulated-tools)
  - [SQLite Memory Backend](#sqlite-memory-backend)
  - [Structured Output Agent](#structured-output-agent)
  - [Tool Timeout & Concurrency](#tool-timeout--concurrency)
  - [Interactive Chat REPL](#interactive-chat-repl)
  - [LangChain Tool Bridge](#langchain-tool-bridge)
- [v0.6 Features](#v06-features)
  - [Argument Auto-Repair](#argument-auto-repair)
  - [Usage Tracking & Budgets](#usage-tracking--budgets)
  - [Loop / Repeated-Call Detection](#loop--repeated-call-detection)
  - [Tool-Name Suggestions](#tool-name-suggestions)
- [Memory](#memory)
- [Observability & Tracing](#observability--tracing)
- [Execution Policies](#execution-policies)
- [MCP Support](#mcp-support)
- [Async & Parallel Tools](#async--parallel-tools)
- [Environment Variables](#environment-variables)
- [Running Tests](#running-tests)
- [Publishing to PyPI](#publishing-to-pypi)

---

## Quick Start

```python
from toolproxy import UniversalAgent, tool

@tool
def search_web(query: str) -> str:
    """Search the web for recent information."""
    return f"Top result for '{query}': ..."

@tool
def calculate(expression: str) -> str:
    """Evaluate a mathematical expression."""
    return str(eval(expression))

agent = UniversalAgent(
    model="openrouter/mistralai/mistral-7b-instruct",
    tools=[search_web, calculate],
)

result = agent.run("What is 15% of 240, and who invented Python?")
print(result.content)
print(f"Steps taken: {result.steps_taken}")
```

---

## How It Works

```
UniversalAgent.run(prompt)
    │
    ├── Planner  ──  auto-detects native vs emulated mode
    │       ├── Native   → provider tool-call API (OpenAI format)
    │       └── Emulated → structured JSON schema injected into system prompt
    │
    ├── Executor  ──  validates args, runs tool, captures errors
    │       └── (optional) ToolCache  →  skip if cached result exists
    │
    └── LoopController  ──  repeats until "final" answer or max_steps
            └── (optional) BaseTracer  →  observability hooks
```

**Emulated mode protocol** — when native tools aren't available, the model is prompted to output exactly one of:

```json
{"type": "tool_call", "tool": {"tool_name": "get_weather", "arguments": {"city": "Chennai"}}}
{"type": "final", "content": "The weather in Chennai is Sunny, 25°C."}
```

Malformed responses are automatically retried (up to `parse_retries` times, default 3) with a corrective error message.

---

## Supported Models & Adapters

| Model Prefix | Backend | Native Tools |
|---|---|---|
| `openrouter/...` | OpenRouter API | ✅ (model-dependent) |
| `ollama/...` | Local Ollama server | ❌ (emulated) |
| `anthropic/...` | Anthropic API | ✅ |
| `gemini/...` | Google Gemini | ✅ |
| `groq/...` | Groq Cloud | ✅ |
| `cohere/...` | Cohere API | ✅ |
| `azure/...` | Azure OpenAI | ✅ |
| *(no prefix)* | OpenAI | ✅ |
| `mock/...` | MockClient — no API key, for testing | configurable |

---

## Core API

### The `@tool` Decorator

Mark any Python function as a tool. Type hints generate the argument schema automatically:

```python
from toolproxy import tool

@tool
def get_stock_price(ticker: str, currency: str = "USD") -> str:
    """Get the current stock price for a ticker symbol."""
    return f"{ticker}: $182.50 {currency}"

# Async tools work too
@tool
async def fetch_data(url: str) -> str:
    """Fetch data from a URL asynchronously."""
    async with httpx.AsyncClient() as client:
        resp = await client.get(url)
        return resp.text
```

For explicit schemas:

```python
from pydantic import BaseModel
from toolproxy import tool

class SearchArgs(BaseModel):
    query: str
    max_results: int = 5

@tool(name="web_search", description="Search the web", args_schema=SearchArgs)
def web_search(query: str, max_results: int = 5) -> str:
    ...
```

### UniversalAgent

```python
from toolproxy import UniversalAgent
from toolproxy.cache import InMemoryToolCache
from toolproxy.memory import ConversationBuffer
from toolproxy.observability import PrintTracer

agent = UniversalAgent(
    model="openrouter/mistralai/mistral-7b-instruct",
    tools=[get_weather, search_web],

    # Mode: "auto" (default), "native_only", "emulated_only"
    mode="auto",

    # Max agent loop iterations
    max_steps=10,

    # API key (or set OPENROUTER_API_KEY env var)
    api_key="sk-...",

    # Custom base URL for OpenAI-compatible endpoints
    base_url="https://my-proxy.example.com/v1",

    # Multi-turn memory (v0.3+)
    memory=ConversationBuffer(),

    # Observability (v0.3+)
    tracer=PrintTracer(),

    # Tool result cache (v0.4+)
    cache=InMemoryToolCache(ttl=300),
)
```

### `run()` and `arun()`

```python
# Synchronous
result = agent.run("What is the weather in Chennai?")

# Asynchronous
result = await agent.arun("What is the weather in Chennai?")

# With options
result = agent.run(
    "Tell me about today's AI news",
    return_trace=True,           # include full tool-call trace
    max_steps=5,                 # override for this run only (v0.4)
    extra_tools=[my_new_tool],   # ad-hoc tools for this run only (v0.4)
    on_tool_call=lambda step, tc: print(f"→ Calling {tc.tool_name}"),
    on_tool_result=lambda step, tr: print(f"← Result: {tr.output}"),
    on_model_output=lambda step, text: print(f"Model: {text}"),
)

print(result.content)       # final answer
print(result.steps_taken)   # number of loop iterations
print(result.trace)         # AgentTrace (if return_trace=True)
```

### Streaming

```python
# Sync streaming
for chunk in agent.stream("Tell me about the solar system"):
    if chunk.tool_call:
        print(f"Calling: {chunk.tool_call.tool_name}")
    elif chunk.tool_result:
        print(f"Result: {chunk.tool_result.output}")
    elif chunk.delta:
        print(chunk.delta, end="", flush=True)
    elif chunk.done:
        print()  # final newline

# Async streaming
async for chunk in agent.astream("Tell me about the solar system"):
    ...
```

---

## v0.4 Features

### Tool Result Caching

Avoid redundant tool calls when the same tool is invoked with identical arguments:

```python
from toolproxy.cache import InMemoryToolCache, FileToolCache

# In-memory: fast, TTL-based, lost on restart
agent = UniversalAgent(
    model="...",
    tools=[search_web],
    cache=InMemoryToolCache(
        ttl=300,        # seconds (None = no expiry)
        max_size=512,   # max entries (LRU eviction)
    ),
)

# File-backed: survives process restarts
agent = UniversalAgent(
    model="...",
    tools=[search_web],
    cache=FileToolCache("tool_cache.json", ttl=3600),
)

# Check cache stats
stats = agent.cache.stats()
# → {"hits": 12, "misses": 3, "size": 8}

# Invalidate a specific entry
agent.cache.invalidate("search_web", {"query": "old topic"})

# Clear all cached results
agent.cache.clear()
```

**How caching works:** The `(tool_name, arguments)` pair is hashed with SHA-256. On a hit, the tool is skipped entirely and the cached string is returned as the tool result. Only successful results are cached — errors are never stored.

---

### Retry with Backoff

Automatically recover from transient failures (network drops, HTTP 429/500/503):

```python
from toolproxy.retry import RetryConfig, with_retry, awith_retry, HTTP_RETRY_CONFIG

# Custom config
config = RetryConfig(
    max_attempts=4,
    initial_delay=1.0,    # seconds before first retry
    backoff_factor=2.0,   # doubles on each retry: 1s, 2s, 4s
    max_delay=30.0,       # capped at 30 seconds
    jitter=True,          # adds ≤10% random jitter
    retryable_exceptions=(ConnectionError, TimeoutError),
)

# Wrap any callable
result = with_retry(lambda: expensive_api_call(), config=config)

# Async version
result = await awith_retry(lambda: async_api_call(), config=config)

# Pre-built config for HTTP model calls
from toolproxy.retry import HTTP_RETRY_CONFIG
result = with_retry(lambda: client.generate(messages), config=HTTP_RETRY_CONFIG)
```

---

### Persistent File Memory

Keep conversation history across process restarts:

```python
from toolproxy.memory import JsonFileMemory

agent = UniversalAgent(
    model="openrouter/mistralai/mistral-7b-instruct",
    tools=[...],
    memory=JsonFileMemory(
        "chat_history.json",
        max_messages=50,        # keep last 50 non-system messages
        preserve_system=True,   # system messages are never trimmed
    ),
)

# Turn 1 (run 1)
agent.run("My name is Alice.")

# Turn 2 (run 2, after restart) — Alice is still remembered
agent.run("What's my name?")
# → "Your name is Alice."
```

Writes are atomic (`.tmp` → rename), so history is never corrupted by a crash.

---

### Agent Extensions

Dynamically modify agents without rebuilding them:

```python
# Add a tool after construction
@tool
def translate(text: str, language: str) -> str:
    """Translate text to a language."""
    ...

agent.add_tool(translate)

# Full reset: clear memory + cache
agent.reset()

# Per-run max_steps override (restores after run)
result = agent.run("complex task", max_steps=25)

# Per-run extra tools (removed after run, even on error)
result = agent.run(
    "use the special tool",
    extra_tools=[one_off_tool],
)

# Check if a tool is registered
"translate" in agent.registry  # → True
len(agent.registry)             # → 3

# Access cache stats
agent.cache.stats()
```

---

### FastAPI Integration

Expose your agent as a production-ready HTTP service:

```python
from fastapi import FastAPI
from toolproxy import UniversalAgent, tool
from toolproxy.integrations.fastapi import create_agent_router

@tool
def search(query: str) -> str:
    """Search for information."""
    return f"Results for {query}..."

agent = UniversalAgent(
    model="openrouter/mistralai/mistral-7b-instruct",
    tools=[search],
)

app = FastAPI(title="My Agent API")
app.include_router(
    create_agent_router(
        agent,
        api_key="my-secret-key",  # optional X-API-Key header auth
    ),
    prefix="/agent",
)
```

**Endpoints created:**

| Method | Path | Description |
|---|---|---|
| `POST` | `/agent/run` | Single-shot: `{"prompt": "..."}` → `{"content": "...", "steps_taken": N}` |
| `POST` | `/agent/stream` | SSE stream: tool events + text deltas + done event |
| `POST` | `/agent/reset` | Clear memory and cache |
| `GET` | `/agent/tools` | List registered tools with descriptions |
| `GET` | `/agent/info` | Model, mode, tool count, native support flag |

**SSE stream events:**
```json
{"type": "tool_call",   "tool": "search", "args": {"query": "AI"}}
{"type": "tool_result", "tool": "search", "output": "AI is ..."}
{"type": "delta",       "text": "Based on "}
{"type": "delta",       "text": "the search..."}
{"type": "done",        "content": "Based on the search, AI is ..."}
```

Test with curl:
```bash
curl -X POST http://localhost:8000/agent/run \
     -H "Content-Type: application/json" \
     -H "X-API-Key: my-secret-key" \
     -d '{"prompt": "What is the weather in Chennai?"}'
```

---

## v0.5 Features

### Parallel Emulated Tools

For models without native function calling support, `toolproxy` emulates tool calling by guiding the model with an advanced structured prompt. Previously, emulated mode was limited to executing one tool at a time. In v0.5, emulated mode can request and execute **multiple tools in parallel** within a single step:

```python
agent = UniversalAgent(
    model="openrouter/google/gemini-2.5-flash",  # emulated tool mode
    tools=[get_weather, get_stock_price],
    allow_parallel_emulated=True,  # enabled by default
)

# If both tools are requested, they execute concurrently in parallel
await agent.arun("What is the weather in Chennai and the stock price of GOOG?")
```

Under the hood, the planner requests a list of JSON action structures in a single prompt response, and the executor dispatches them concurrently.

---

### SQLite Memory Backend

`SqliteMemory` is a persistent memory provider backed by a local SQLite database file. It enables multi-session isolation, keyword searching, and history trimming.

```python
from toolproxy.memory import SqliteMemory

memory = SqliteMemory(
    db_path="chat_history.db",
    session_id="session-user-42",
    max_messages=20,       # trim older messages when limit is reached
    preserve_system=True,  # keep the system prompt intact when trimming
)

agent = UniversalAgent(
    model="...",
    memory=memory,
)
```

**Additional Features:**
- **Search:** Query the database for past messages using SQL `LIKE` wildcard matching.
  ```python
  # Search within the current session:
  results = memory.search("Python")
  
  # Search across all sessions by passing session_id="":
  all_results = memory.search("Python", session_id="")
  ```
- **Sessions:** Get all active session IDs stored in the database.
  ```python
  active_sessions = memory.sessions()
  ```
- **Message Count:** Query the number of messages.
  ```python
  count = memory.message_count()  # or len(memory)
  ```

---

### Structured Output Agent

`StructuredOutputAgent` wraps a `UniversalAgent` and ensures that the final response validates against a specific Pydantic model instead of a raw text string. It automatically appends the JSON schema requirements to the prompt, validates the model output, and performs an automatic format-retry loop if validation fails.

```python
from pydantic import BaseModel
from toolproxy import UniversalAgent
from toolproxy.structured import StructuredOutputAgent

class MovieDetails(BaseModel):
    title: str
    release_year: int
    genres: list[str]
    rating: float

agent = UniversalAgent(model="openai/gpt-4o")
structured_agent = StructuredOutputAgent(agent, output_schema=MovieDetails, max_retries=1)

response = structured_agent.run("Tell me about Inception")
# response is a StructuredAgentResponse containing the validated model instance
movie = response.data
print(f"{movie.title} ({movie.release_year}) - Rating: {movie.rating}")
```

---

### Tool Timeout & Concurrency

Ensure slow external tools don't freeze your agent, and limit concurrency to prevent rate-limiting or overloading external APIs.

```python
agent = UniversalAgent(
    model="...",
    tools=[slow_tool, other_tool],
    tool_timeout=5.0,     # cancel any tool execution exceeding 5.0 seconds
    max_concurrency=3,    # run at most 3 tools in parallel concurrently
)
```

- When a tool exceeds the timeout, `toolproxy` raises a `ToolTimeoutError` and logs the event, reporting the failure back to the agent so it can adapt or retry.

---

### Interactive Chat REPL

Quickly interact with your agent in a command-line interface. It handles multi-turn conversation memory, streams response tokens as they arrive, and prints visual indicators when tools are called.

```python
agent = UniversalAgent(model="openai/gpt-4o", tools=[get_weather])

# Start the REPL
agent.chat(
    welcome_msg="Welcome to the Interactive Agent Chat! Type 'quit' to exit.",
    stream=True,  # stream tokens as they arrive
)
```

---

### LangChain Tool Bridge

The LangChain bridge allows you to reuse LangChain tools within `toolproxy` and vice-versa, allowing for easy migration and interoperability.

```python
from toolproxy.integrations.langchain import from_langchain_tool, to_langchain_tool

# 1. Use a LangChain tool in toolproxy:
from langchain_community.tools import WikipediaQueryRun
wikipedia = WikipediaQueryRun(...)

tp_tool = from_langchain_tool(wikipedia)
agent = UniversalAgent(model="...", tools=[tp_tool])

# 2. Use a toolproxy tool in LangChain:
from toolproxy import tool

@tool
def custom_calc(a: int, b: int) -> int:
    "Add two numbers."
    return a + b

lc_tool = to_langchain_tool(custom_calc)
# Now pass lc_tool to your LangChain agent's tools list!
```

---

## v0.6 Features

**Reliability & Guardrails.** The most common ways LLM tool calling fails in
production are well documented: models emit *malformed or misspelled arguments*,
get *stuck calling the same tool in a loop*, and *quietly run up cost* on a
runaway agent. v0.6 ships four guardrails — all opt-in, all provider-agnostic —
that defend against exactly those failure modes.

```python
agent = UniversalAgent(
    model="openrouter/mistralai/mistral-7b-instruct",
    tools=[get_weather],
    repair_arguments=True,    # fix almost-right tool args before they fail
    max_tool_calls=20,        # hard cap on tool executions per run
    token_budget=8000,        # hard cap on approx tokens per run
    max_repeated_calls=3,     # abort if the same call repeats 3×
)

result = agent.run("What's the weather in Chennai?")
print(result.usage)   # → Usage(steps=2, tool_calls=1, tokens=143)
```

### Argument Auto-Repair

Models frequently produce arguments that are *almost* right — a misspelled
parameter name, a number sent as a string, a JSON object sent as a string, or
the real arguments wrapped in an extra `{"arguments": {...}}` envelope. With
`repair_arguments=True`, these are coerced to match the tool's schema **before**
validation, so the agent makes progress instead of bouncing an error back:

```python
agent = UniversalAgent(model="...", tools=[get_weather], repair_arguments=True)

# Model emits {"citi": "Paris", "days": "3"} →
# repaired to {"city": "Paris", "days": 3} and the tool runs successfully.
```

Repairs applied (each is conservative and only fires when unambiguous):

| Problem | Example | Repaired to |
|---|---|---|
| Misspelled key (fuzzy match) | `{"citi": "Paris"}` | `{"city": "Paris"}` |
| Stringified number | `{"days": "3"}` | `{"days": 3}` |
| Stringified bool | `{"metric": "no"}` | `{"metric": false}` |
| JSON string for object/array | `{"opts": '{"a": 1}'}` | `{"opts": {"a": 1}}` |
| Wrapped payload | `{"arguments": {...}}` | `{...}` |

The repair is also exposed standalone:

```python
from toolproxy import repair_arguments
from pydantic import BaseModel

class Args(BaseModel):
    city: str
    days: int

repair_arguments(Args, {"citi": "Paris", "days": "3"})
# → {"city": "Paris", "days": 3}
```

Anything it cannot confidently fix is passed through untouched, so normal
Pydantic validation still produces the canonical error message.

### Usage Tracking & Budgets

Every run now returns a `Usage` accounting object, and you can enforce hard
ceilings to guard against runaway cost and latency:

```python
agent = UniversalAgent(
    model="...",
    tools=[search],
    max_tool_calls=20,   # raise BudgetExceededError after 20 tool calls
    token_budget=8000,   # raise BudgetExceededError past ~8000 tokens
)

result = agent.run("Research this topic thoroughly")
print(result.usage.steps)       # loop iterations
print(result.usage.tool_calls)  # tools executed
print(result.usage.tokens)      # approximate tokens (chars/4 heuristic)
```

When a limit is exceeded mid-run, `BudgetExceededError` is raised with the
`limit_kind` (`"tool_calls"` or `"tokens"`), the `limit`, and the amount `used`.
Limits are **per-run**, not for the agent's lifetime.

### Loop / Repeated-Call Detection

A classic failure is the model calling the same tool with identical arguments
over and over. `max_repeated_calls` breaks the loop early:

```python
agent = UniversalAgent(model="...", tools=[get_weather], max_repeated_calls=3)

# If the model requests get_weather(city="Paris") a 3rd identical time,
# RepeatedToolCallError is raised instead of burning the whole step budget.
```

### Tool-Name Suggestions

When a model invents or misspells a tool name, `ToolNotFoundError` now suggests
the closest real tool, and that hint is fed back into the conversation so the
model can self-correct:

```
ToolNotFoundError: Tool 'get_wether' not found in registry. Did you mean 'get_weather'?
```

---

## Memory

Keep context across multiple `agent.run()` calls:

```python
from toolproxy.memory import (
    ConversationBuffer,    # unlimited history
    SlidingWindowMemory,   # last N messages
    TokenWindowMemory,     # fits within token budget
    JsonFileMemory,        # persistent across restarts (v0.4)
)

# Unlimited
memory = ConversationBuffer()

# Last 20 messages (keeps system messages)
memory = SlidingWindowMemory(max_messages=20)

# Within 4096 tokens (uses char/4 approximation)
memory = TokenWindowMemory(max_tokens=4096)

# Persistent JSON file
memory = JsonFileMemory("history.json", max_messages=50)

agent = UniversalAgent(model="...", tools=[...], memory=memory)

agent.run("My name is Alice.")
agent.run("What's my name?")   # remembers: "Your name is Alice."

agent.clear_memory()           # wipe history without rebuilding agent
```

---

## Observability & Tracing

```python
from toolproxy.observability import FileTracer, PrintTracer, OpenTelemetryTracer

# Print to console
agent = UniversalAgent(..., tracer=PrintTracer(prefix="[agent]"))

# Write structured JSON to a file
agent = UniversalAgent(..., tracer=FileTracer("runs.jsonl", mode="append"))

# Emit OpenTelemetry spans
from opentelemetry import trace
tracer_provider = trace.get_tracer_provider()
agent = UniversalAgent(..., tracer=OpenTelemetryTracer(tracer_provider))

# Read recorded runs from a FileTracer
tracer = FileTracer("runs.jsonl", mode="append")
runs = tracer.read_runs()
for run in runs:
    print(run["prompt"], run["steps"])
```

---

## Execution Policies

Control which tools the model is allowed to call:

```python
from toolproxy.config import ExecutionPolicy

# Allow all tools (default)
policy = ExecutionPolicy(mode="allow_all")

# Whitelist specific tools
policy = ExecutionPolicy(mode="allow_only", allowed_tools=["get_weather"])

# Require confirmation before dangerous tools
policy = ExecutionPolicy(
    mode="confirm_before",
    confirm_tools=["delete_file", "send_email"],
)

agent = UniversalAgent(
    model="...",
    tools=[...],
    execution_policy=policy,
    confirm_callback=lambda tool_name, args: input(f"Allow {tool_name}({args})? [y/N] ") == "y",
)
```

---

## MCP Support

Connect to any [Model Context Protocol](https://modelcontextprotocol.io) server and use its tools:

```python
import asyncio
from toolproxy import UniversalAgent
from toolproxy.mcp import MCPToolset

async def main():
    # Discover tools from an MCP server
    toolset = await MCPToolset.from_server("http://localhost:3000")
    print(f"Discovered {len(toolset)} tools: {[t.name for t in toolset.tools]}")

    agent = UniversalAgent(
        model="openrouter/mistralai/mistral-7b-instruct",
    )
    toolset.register_into(agent.registry)

    result = await agent.arun("Use the MCP tools to answer my question.")
    print(result.content)

asyncio.run(main())
```

---

## Async & Parallel Tools

toolproxy fully supports async — including parallel tool dispatch:

```python
import asyncio
from toolproxy import UniversalAgent, tool

@tool
async def fetch_page(url: str) -> str:
    """Fetch a webpage asynchronously."""
    async with httpx.AsyncClient() as c:
        return (await c.get(url)).text

agent = UniversalAgent(
    model="openrouter/mistralai/mistral-7b-instruct",
    tools=[fetch_page],
)

# When the model requests multiple tools at once (native mode),
# they are dispatched concurrently via asyncio.gather
result = await agent.arun("Fetch the homepages of openai.com and anthropic.com")
```

---

## Environment Variables

| Variable | Description | Default |
|---|---|---|
| `OPENROUTER_API_KEY` | API key for OpenRouter | — |
| `OPENAI_API_KEY` | API key for OpenAI | — |
| `ANTHROPIC_API_KEY` | API key for Anthropic | — |
| `GEMINI_API_KEY` | API key for Google Gemini | — |
| `GROQ_API_KEY` | API key for Groq | — |
| `COHERE_API_KEY` | API key for Cohere | — |
| `AZURE_OPENAI_API_KEY` | API key for Azure OpenAI | — |
| `AZURE_OPENAI_ENDPOINT` | Azure OpenAI endpoint URL | — |
| `OLLAMA_BASE_URL` | Ollama server URL | `http://localhost:11434` |
| `OLLAMA_MODEL` | Ollama model name | `llama3` |

---

## Project Structure

```
src/toolproxy/
  __init__.py          # Public API -- all exports
  agent.py             # UniversalAgent (+ chat() REPL v0.5)
  llm_client.py        # LLMClient + provider adapters
  tools.py             # @tool decorator + ToolRegistry
  schemas.py           # Pydantic schemas (Message, ToolCall, Action, ...)
  planner.py           # Prompt building + response parsing (parallel emulated v0.5)
  executor.py          # Tool execution + policy + cache + timeout/concurrency [v0.5]
  loop.py              # LoopController (orchestration, parallel emulated v0.5)
  cache.py             # InMemoryToolCache, FileToolCache      [v0.4]
  retry.py             # RetryConfig, with_retry, awith_retry  [v0.4]
  guardrails.py        # repair_arguments, UsageTracker,
                       # CallSignatureCounter                  [v0.6]
  memory.py            # ConversationBuffer, SlidingWindow, TokenWindow,
                       # JsonFileMemory [v0.4], SqliteMemory [v0.5]
  structured.py        # StructuredOutputAgent                 [v0.5]
  config.py            # AgentConfig, ExecutionPolicy, capability map
  exceptions.py        # Typed exception hierarchy
  mcp.py               # MCPClient, MCPToolset                 [v0.3]
  observability/       # BaseTracer, FileTracer, PrintTracer, OTelTracer
  integrations/
    fastapi.py         # create_agent_router()                 [v0.4]
    langchain.py       # from/to_langchain_tool()              [v0.5]

examples/
  basic_chat.py
  openrouter_tools.py
  local_ollama.py
  caching_demo.py      # [v0.4]
  persistent_memory.py # [v0.4]
  fastapi_agent.py     # [v0.4]
  parallel_tools.py    # [v0.5]
  structured_output.py # [v0.5]
  chat_repl.py         # [v0.5]
  argument_repair_demo.py    # [v0.6]
  budget_guardrails_demo.py  # [v0.6]
  loop_detection_demo.py     # [v0.6]

tests/
  test_agent_basic.py
  test_emulated_mode.py
  test_native_mode.py
  test_error_handling.py
  test_tool_registry.py
  test_memory.py
  test_observability.py
  test_mcp.py
  test_cache.py        # [v0.4]
  test_retry.py        # [v0.4]
  test_file_memory.py  # [v0.4]
  test_agent_extensions.py # [v0.4]
  test_fastapi.py      # [v0.4]
  test_parallel_emulated.py # [v0.5]
  test_sqlite_memory.py     # [v0.5]
  test_structured.py        # [v0.5]
  test_executor_timeout.py  # [v0.5]
  test_chat_repl.py         # [v0.5]
  test_langchain_bridge.py  # [v0.5]
  test_guardrails.py        # [v0.6]
```

---

## Running Tests

```bash
# Run all tests
pytest tests/ -v

# Run a specific test file
pytest tests/test_cache.py -v

# Run with coverage
pytest tests/ --cov=toolproxy --cov-report=term-missing
```

All tests use `MockClient` — **no API keys required**.

---

## Optional Dependencies

| Extra | Installs | Enables |
|---|---|---|
| `pip install toolproxy[anthropic]` | `anthropic` | Anthropic Claude adapter |
| `pip install toolproxy[fastapi]` | `fastapi`, `httpx` | FastAPI router integration |
| `pip install toolproxy[otel]` | `opentelemetry-sdk` | OpenTelemetry tracing |
| `pip install toolproxy[full]` | All of the above | Everything |

---

## Publishing to PyPI

```bash
# 1. Install build tools
pip install build twine

# 2. Build wheel + sdist
python -m build

# 3. Verify the distribution
twine check dist/*

# 4. Upload to TestPyPI first (recommended)
twine upload --repository testpypi dist/*

# 5. Upload to PyPI
twine upload dist/*
```

---

## Changelog

### v0.6.0 — Reliability & Guardrails
- 🆕 **Argument auto-repair** — `repair_arguments=True` fixes the most common tool-format hallucinations before validation: fuzzy parameter-name matching, JSON-string unwrapping, scalar coercion, and `{"arguments": {...}}` payload unwrapping. Also exposed standalone via `from toolproxy import repair_arguments`.
- 🆕 **Usage tracking & budgets** — every run returns `result.usage` (`steps`, `tool_calls`, `tokens`); `max_tool_calls` and `token_budget` raise `BudgetExceededError` to stop runaway cost/latency.
- 🆕 **Loop / repeated-call detection** — `max_repeated_calls` raises `RepeatedToolCallError` when the model is stuck calling the same tool with identical arguments.
- 🆕 **Tool-name suggestions** — `ToolNotFoundError` now appends a "Did you mean '…'?" hint for the closest registered tool.
- 🆕 New exports: `repair_arguments`, `UsageTracker`, `CallSignatureCounter`, `approx_tokens`, `Usage`, `BudgetExceededError`, `RepeatedToolCallError`.
- 🆕 **Docs site** — a static, no-build GitHub Pages site under `docs/`.
- 📈 Test suite: 369 → 396 passing tests

### v0.5.0
- 🆕 **Parallel emulated tools** — non-native models can now call multiple tools per step (`"type": "tool_calls"` action)
- 🆕 **SQLite memory** — `SqliteMemory` with multi-session support, `search()`, `sessions()`, `message_count()`
- 🆕 **StructuredOutputAgent** — get a typed Pydantic model back instead of a string
- 🆕 **Tool timeout** — `tool_timeout=N` raises `ToolTimeoutError` for slow tools (sync + async)
- 🆕 **Max concurrency** — `max_concurrency=N` caps parallel async tool calls via `asyncio.Semaphore`
- 🆕 **`agent.chat()`** — one-line interactive terminal REPL with tool indicators
- 🆕 **LangChain bridge** — `from_langchain_tool()` / `to_langchain_tool()` (install: `pip install toolproxy[langchain]`)
- 📈 Test suite: 274 → 369 passing tests

### v0.4.0
- 🆕 **Tool result caching** — `InMemoryToolCache`, `FileToolCache` with TTL and LRU eviction
- 🆕 **Retry with exponential backoff** — `RetryConfig`, `with_retry`, `awith_retry`, `HTTP_RETRY_CONFIG`
- 🆕 **Persistent memory** — `JsonFileMemory` with atomic writes and trimming
- 🆕 **Agent extensions** — `add_tool()`, `reset()`, per-run `max_steps` and `extra_tools`
- 🆕 **FastAPI integration** — `create_agent_router()` with SSE streaming and API key auth
- 📈 Test suite: 172 → 274 passing tests

### v0.3.0
- Multi-turn conversation memory (`ConversationBuffer`, `SlidingWindowMemory`, `TokenWindowMemory`)
- Structured observability (`FileTracer`, `PrintTracer`, `OpenTelemetryTracer`)
- Model Context Protocol (MCP) client and toolset integration
- New adapters: Anthropic, Gemini, Groq, Cohere, Azure OpenAI
- Async streaming (`agent.stream()`, `agent.astream()`)

### v0.2.0
- Async support (`arun()`, async tool execution with `asyncio.gather`)
- Execution policies (`allow_all`, `allow_only`, `confirm_before`)
- Streaming events (`StreamChunk`)
- `probe_tool_support()` async utility

### v0.1.0
- Initial release: `UniversalAgent`, `@tool`, emulated + native modes
- OpenRouter, OpenAI, Ollama backends
- `ConversationBuffer` memory, `AgentTrace`

---

## License

MIT — see [LICENSE](LICENSE).
