Metadata-Version: 2.4
Name: agent-forge-svkmsr6
Version: 0.1.0
Summary: A production-shaped multi-agent orchestration skeleton on LangGraph with first-class tracing, retries, and token budgets.
Author-email: Souvik Misra <svkmsr6@gmail.com>
License: MIT
Project-URL: Repository, https://github.com/svkmsr6/agent-forge
Keywords: agents,langgraph,orchestration,observability,llm
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: langgraph>=0.2.0
Requires-Dist: langchain-core>=0.2.0
Requires-Dist: openai>=1.30.0
Requires-Dist: opentelemetry-sdk>=1.25.0
Requires-Dist: pydantic>=2.6
Requires-Dist: tenacity>=8.2
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: ruff>=0.5.0; extra == "dev"
Requires-Dist: mypy>=1.10; extra == "dev"
Dynamic: license-file

# agent-forge

A production-shaped multi-agent orchestration skeleton on LangGraph — with tracing, retry/fallback policies, and token budgets treated as first-class concerns, not afterthoughts.

I built this because most agent demos I reviewed while designing production systems stopped at "chain some prompts together." The hard 20% — observability, cost control, failure policy — was always left as an exercise for the reader. This repo is my opinionated take on what that missing 20% looks like, extracted into a small, reusable framework skeleton.

## Architecture

```mermaid
flowchart LR
    Q[User query] --> AF[AgentForge.run]
    AF --> GB[GraphBuilder]
    GB --> LG[(LangGraph StateGraph)]
    LG --> N1[plan node]
    N1 --> N2[execute node]
    N2 --> N3[synthesize node]
    N2 --> TR[ToolRegistry]
    TR --> T1[CalculatorTool]
    TR --> T2[WebSearchTool stub]
    N1 -. wrapped .-> OBS[OTel tracing]
    N2 -. wrapped .-> OBS
    N3 -. wrapped .-> OBS
    N1 -. metered .-> BUD[BudgetTracker]
    N2 -. metered .-> BUD
    N3 -. metered .-> BUD
    N1 & N2 & N3 -. retried .-> POL[Retry / fallback policies]
```

Every node that enters the graph passes through three cross-cutting wrappers, applied in this order:

1. **Tracing** — an OpenTelemetry span is opened per node execution (name, duration, token usage as span attributes). Export is your choice: OTel Collector, Jaeger, or LangSmith via env vars.
2. **Budget** — token and cost usage recorded against a per-run `RunBudget`. Exceeding the budget raises `BudgetExceeded`, which halts the run cleanly instead of producing a surprise invoice.
3. **Policy** — nodes that call an LLM go through a tenacity-based retry decorator; model calls go through a fallback chain (`primary → secondary → ...`) so a provider outage degrades instead of failing.

## Features

- **Graph builder with guard rails** — register typed nodes, wire edges, add conditional routing; validates the graph (dangling edges, missing entry point) before LangGraph ever sees it.
- **Tool plugin interface** — one `BaseTool` contract (`name`, `description`, `run()`); ships with a safe-AST calculator and an honestly-stubbed web-search tool with a clear extension point.
- **Retry & fallback policies** — a tenacity-based decorator factory, plus a `FallbackChain` that routes model calls across providers.
- **Per-run token/cost budgets** — a `BudgetTracker` with aggregate and per-node caps, scoped per run via a `contextvars` token so concurrent runs don't leak usage into each other.
- **OpenTelemetry tracing** — spans around every node execution; optional LangSmith-compatible env configuration.
- **Deterministic mock LLM** — the example runs end-to-end with no API key, so CI and first-time contributors get a green run out of the box.

## Quickstart (60 seconds)

```bash
pip install -e ".[dev]"
python examples/research_agent.py
```

Expected output (abridged):

```
== agent-forge example: research agent ==
Query: What is the capital of France, and what is 17 * 24?

--- plan ---
[1] answer-geography: Identify the capital of France.
[2] answer-math: Compute 17 * 24.

--- execute ---
answer-math -> 408
...

--- synthesize ---
The web search tool is not configured in this demo environment. 408.
Usage: input=300 output=150 total=450
Nodes visited: plan -> execute -> synthesize
```

Point it at a real model by setting `OPENAI_API_KEY` and swapping the mock client in `examples/research_agent.py` — the extension point is marked with a comment.

## Tech stack

| Layer | Choice | Why |
| --- | --- | --- |
| Orchestration | LangGraph | Stateful, cyclical graphs; checkpointing story; not a prompt-chaining framework |
| Data contracts | Pydantic v2 | Runtime-validated state, budgets, and tool results |
| Resilience | Tenacity | Battle-tested retry semantics (wait, stop, retry-on) |
| Observability | OpenTelemetry SDK | Vendor-neutral traces; LangSmith compatible via env config |
| LLM client | OpenAI SDK | Kept behind a thin client protocol so other providers slot in |
| Tooling | Ruff, mypy (strict), pytest | Fast lint + strict typing; tests mock all LLM boundaries |

## Design decisions

The rationale behind the big calls lives in [`docs/adr/`](docs/adr). The short version:

- **LangGraph over CrewAI** (ADR-0001): I need control over state shape, routing, and checkpoints — CrewAI's role abstraction hides exactly the seams I care about.
- **OpenTelemetry for tracing** (ADR-0002): agent traces are most valuable next to your existing infra traces, not in a separate silo. OTel spans carry token usage as first-class attributes.
- **Budget as a first-class concern** (ADR-0003): cost control is a cross-cutting runtime policy (like auth), not something each node should remember to do. The tracker lives in a context var; nodes report usage, the framework enforces limits.

Full system design write-ups: [HLD](docs/hld.md) and [LLD](docs/lld.md).

## Repository layout

```
src/agent_forge/
  agent.py       # AgentForge: public facade tying everything together
  graph.py       # GraphBuilder, ForgeState, node wrapper (budget + tracing + policy)
  budget.py      # RunBudget, BudgetTracker, BudgetExceeded
  policies.py    # retry decorator factory, FallbackChain
  tools/         # BaseTool, ToolRegistry, calculator + web-search stub
  tracing.py     # OTel span wrappers, LangSmith env config
examples/
  research_agent.py   # plan -> execute -> synthesize, mock LLM, no API key needed
docs/
  hld.md  lld.md  adr/
tests/
```

## Roadmap

- [ ] Checkpointing / durable execution via LangGraph checkpointers
- [ ] Streaming intermediate node outputs to the caller
- [ ] Cost tables per model (tokens → USD) loaded from config
- [ ] Human-in-the-loop interrupt nodes (approval gates)
- [ ] Async node execution for I/O-bound tool calls
- [ ] A second example: tool-using customer-support agent with fallback model routing

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md). PRs welcome; CI runs ruff, mypy, and pytest on Python 3.10–3.12.

## License

MIT — see [LICENSE](LICENSE).
