Metadata-Version: 2.4
Name: enterprise-agentic-ai-framework
Version: 0.2.0
Summary: Enterprise Agentic AI Framework SDK
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: >=3.10
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic>=2.0
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == 'dev'
Description-Content-Type: text/markdown

# enterprise-agentic-ai-framework

An enterprise governance framework for building single- and multi-agent
AI systems in Python: authorization, guardrails, observability, secrets
management, LLM gateway access, and a full production evaluation
suite, all as one consistent stack instead of one-off code per project.

```bash
pip install enterprise-agentic-ai-framework
```

The import name is `agentic_ai` (the PyPI distribution name is longer
for naming reasons, the package you actually `import` is not):

```python
from agentic_ai.gateway import LiteLLMGateway
```

## Status

This is an early release. **The LLM gateway and the full evaluation
suite are implemented today** - everything else below is scaffolded
(the module exists, it's empty) and not yet usable. This table will be
kept current as modules land, not written once and left stale.

| Module | Status |
|---|:---:|
| `gateway` - LLM gateway (LiteLLM proxy client) | ✅ Implemented |
| `evaluation` - Agent/LLM/Tools/Multi-Agent/RAG/Security/Platform/Memory/Drift evaluation (48 metrics, see below) | ✅ Implemented |
| `identity` - authentication | ⏳ Planned |
| `governance` - authorization (PEP/PDP) | ⏳ Planned |
| `guardrails` - PII/secrets/injection/jailbreak detection | ⏳ Planned |
| `secrets` - secrets management | ⏳ Planned |
| `observability` - distributed tracing, structured audit | ⏳ Planned |
| `memory` - short/long-term agent memory (storage & retrieval itself) | ⏳ Planned |
| `context` - context engineering (write/select/compress) | ⏳ Planned |
| `finops` - LLM cost tracking | ⏳ Planned |
| `security` - rate limiting, abuse detection (live enforcement) | ⏳ Planned |
| `compliance`, `audit`, `data_governance` | ⏳ Planned |
| `monitoring`, `resilience`, `responsible_ai` | ⏳ Planned |
| `core` - agent/tool base classes, orchestrator | ⏳ Planned |

**A naming note, not a contradiction**: `evaluation.memory` and
`evaluation.security` are implemented; the top-level `memory` and
`security` modules are not. They're different things - `evaluation.*`
*measures* something (was a memory retrieval correct? did PII leak in
a response you already captured?) from data you already collected,
which needs no live enforcement layer underneath it. The top-level
`memory`/`security` modules would *be* that live layer (actually
storing conversation memory, actually rate-limiting requests) - planned,
not built yet.

## Prerequisites

**This library is a client, not a server.** Before any of the examples
below will work, you need a LiteLLM proxy already running somewhere
reachable - `agentic_ai.gateway` never installs, starts, stops, or
otherwise manages that process for you. Set it up once:

**1. Install LiteLLM's proxy** (a separate package from this library):

```bash
pip install 'litellm[proxy]'
```

**2. Register at least one model.** Create `litellm_config.yaml` -
this example routes the model name `gpt-4o-mini` to OpenAI, reading the
real provider key from an environment variable (never hardcode it in
the YAML):

```yaml
model_list:
  - model_name: gpt-4o-mini
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY
```

Any provider LiteLLM supports works the same way - Anthropic, Azure
OpenAI, Bedrock, a local Ollama model, etc.; only `litellm_params`
changes. See LiteLLM's own docs for the full provider list.

**3. Set the real provider key and start the proxy:**

```bash
export OPENAI_API_KEY=sk-...
litellm --config litellm_config.yaml --port 4000
```

**4. Confirm it's actually up** before writing any Python against it:

```bash
curl http://localhost:4000/health/liveliness
# -> "I'm alive!"
```

If that curl fails, nothing below will work either - fix connectivity
to the proxy first; `agentic_ai.gateway`'s errors will otherwise (correctly)
just tell you the same thing: it can't reach `http://localhost:4000`.

Only once you have a real, running, reachable LiteLLM proxy do the
examples below have anything to talk to.

## Quickstart: LLM Gateway

### 1. Connect to it

```python
from agentic_ai.gateway import LiteLLMGateway

# No arguments needed for the common case: connects to
# http://localhost:4000, LiteLLM's own default port.
gateway = LiteLLMGateway()

reply = gateway.complete(
    model="gpt-4o-mini",  # must be registered on your proxy, e.g. in litellm_config.yaml
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Name three benefits of distributed tracing."},
    ],
)
print(reply)
```

### 2. Configuring host, port, and auth

```python
from agentic_ai.gateway import LiteLLMGateway

# Custom port - your proxy isn't on LiteLLM's default 4000
gateway = LiteLLMGateway(port=5001)

# Custom host and port - a proxy running elsewhere on your network
gateway = LiteLLMGateway(host="litellm.internal", port=8080)

# Full base_url - anything host/port can't express (TLS, a path prefix)
gateway = LiteLLMGateway(base_url="https://litellm.example.com/proxy")

# A proxy that requires a virtual key
gateway = LiteLLMGateway(api_key="sk-...")  # resolve this from your own
                                             # secrets store - the gateway
                                             # module doesn't fetch it for you
```

### 3. The full response, not just the text

`complete()` is a convenience wrapper around `chat_completion()`, which
returns the full OpenAI-compatible response body (usage, finish_reason,
etc.) when you need more than just the message content:

```python
result = gateway.chat_completion(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Summarize this in one sentence: ..."}],
    temperature=0.2,
    max_tokens=200,
)
print(result["choices"][0]["message"]["content"])
print(result["usage"])
```

### 4. Handling errors

The gateway never lets a raw network exception escape - callers get one
of two exceptions, so "the proxy is down" and "the proxy rejected the
request" are never conflated:

```python
from agentic_ai.gateway import GatewayConnectionError, GatewayRequestError, LiteLLMGateway

gateway = LiteLLMGateway()

try:
    reply = gateway.complete("gpt-4o-mini", [{"role": "user", "content": "hi"}])
except GatewayConnectionError:
    # Nothing is listening at gateway.base_url at all - is LiteLLM
    # actually running? (see Prerequisites above)
    ...
except GatewayRequestError as e:
    # The proxy responded, but with an error (bad model name, missing
    # api_key, malformed request) - e includes the proxy's own message.
    print(e)
```

### 5. Cleaning up

`LiteLLMGateway` holds an open HTTP connection pool; close it when
you're done, or use it as a context manager:

```python
with LiteLLMGateway() as gateway:
    reply = gateway.complete("gpt-4o-mini", [{"role": "user", "content": "hi"}])
# connection pool closed automatically here
```

## Evaluation

A complete production evaluation surface for single- and multi-agent AI
systems - 48 metrics across 9 categories, organized one folder per
category under `agentic_ai.evaluation`:

| Category | Import | Measures |
|---|---|---|
| Agent | `agentic_ai.evaluation.agent` | Task Success/Correctness, Planning, Reasoning, Execution, Recovery, Autonomy, Loops, Lifecycle |
| LLM | `agentic_ai.evaluation.llm` | Response Correctness, Groundedness, Hallucination Rate, Instruction Following, Safety/Policy Violation, Latency/Tokens/Cost |
| Tools | `agentic_ai.evaluation.tools` | Selection/Argument Accuracy, Success/Failure Rate, Unnecessary Calls, Latency |
| Multi-Agent | `agentic_ai.evaluation.multi_agent` | Routing, Delegation, Handoff, Coordination, Duplicate Work |
| RAG | `agentic_ai.evaluation.rag` | Recall@K, Context Relevance, Groundedness, Citation Accuracy |
| Security | `agentic_ai.evaluation.security` | Prompt Injection, Unauthorized Execution, PII/Cross-Tenant Leakage, Authorization Violations |
| Platform | `agentic_ai.evaluation.platform` | Error Rate, Timeout Rate, Cost per Successful Task, SLA Compliance |
| Memory | `agentic_ai.evaluation.memory` | Retrieval Accuracy, Consistency |
| Drift | `agentic_ai.evaluation.drift` | Statistical (z-score) drift on Success/Correctness/Hallucination/Latency/Cost |

Every category is deterministic, LLM-judged, or a documented mix of
both - deterministic metrics need no LLM call at all (they read fields
you already populated); judged metrics reuse the same `LLMJudge` from
`agentic_ai.evaluation.core`, built on the gateway above, nothing else.

### Deterministic - no LLM call needed

```python
from agentic_ai.evaluation.agent import AgentTrace, compute_task_execution

traces = [
    AgentTrace(run_id="r1", task="find backend jobs", task_succeeded=True),
    AgentTrace(run_id="r2", task="find backend jobs", task_succeeded=False),
    AgentTrace(run_id="r3", task="find backend jobs", task_succeeded=True),
]
metrics = compute_task_execution(traces)
print(metrics.success_rate)  # 0.6666666666666666
```

### LLM-judged - needs a gateway, same one as above

```python
from agentic_ai.evaluation import LLMJudge
from agentic_ai.evaluation.llm import LLMCall, judge_response_correctness
from agentic_ai.gateway import LiteLLMGateway

judge = LLMJudge(LiteLLMGateway(), model="gpt-4o-mini")
call = LLMCall(call_id="c1", model="gpt-4o-mini", prompt="What is 2+2?", response="4")

result = judge_response_correctness(judge, call)
print(result.correct, result.score)
```

Every `judge_*()` function across every category takes an optional
`system_prompt` override - the built-in `DEFAULT_*` rubric is a real,
usable starting point, not the only valid one for every domain:

```python
from agentic_ai.evaluation.llm import judge_response_correctness

legal_rubric = "You are a strict legal-domain correctness judge. ..."
result = judge_response_correctness(judge, call, system_prompt=legal_rubric)
```

### Everything at once, persisted, compared over time

Agent Evaluation ties every deterministic + judged category together
into one report, storable and diffable:

```python
from agentic_ai.evaluation.agent import evaluate, JSONLEvaluationStore, compare

report = evaluate(traces, judge=judge)  # runs every computable category
store = JSONLEvaluationStore("eval_runs.jsonl")
store.save(report)

baseline = store.list_runs(limit=2)[1]
regressions = compare(baseline, report)  # direction-aware: knows failure_rate up is bad
```

For statistical drift across many runs over time (not just two points),
see `agentic_ai.evaluation.drift.compute_drift()` and its five named
wrappers (`compute_task_success_drift`, `compute_correctness_drift`,
`compute_hallucination_drift`, `compute_latency_drift`,
`compute_cost_drift`).

### Every category's own trace/call shape

`agent`, `llm`, `multi_agent`, `rag`, `security`, and `memory` each have
their own input model (`AgentTrace`, `LLMCall`, `MultiAgentTrace`,
`RAGQuery`, `AuthorizationCheck`/`TenantDataCheck`,
`MemoryRetrieval`) - populate the one your category needs from your own
agent's logging; nothing in this library runs an agent or a retriever
for you, it only evaluates the record you hand it.

## Requirements

- Python 3.10+
- A LiteLLM proxy you deploy yourself (this library is a client, not a
  bundled server)

## License

Apache-2.0
