Metadata-Version: 2.4
Name: route66
Version: 1.1.0
Summary: route66 Python SDK
Author: route66 maintainers
License: Proprietary
Keywords: route66,sdk
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Requires-Dist: twine>=5.1; extra == 'dev'
Provides-Extra: semantic
Requires-Dist: sentence-transformers>=3.0; extra == 'semantic'
Description-Content-Type: text/markdown

# route66

**Intelligent Tool Router for LLM Agents** — a catalog-in library that,
given a query and a large tool catalog, returns only the relevant tools,
deduplicated, in an ordered execution plan, with policy-aware defaults.

Built against the **`router-testkit`** hackathon grading contract
(single-tool routing, ordered multi-tool plans, near-duplicate
resolution, capability-based fallback widening, deprecation policy,
out-of-scope refusal, and clarify-on-vague-destructive).

## Layout

```
route66/
├── tool_router/          # The POC library (facade + retrievers + dedup)
├── testkit_adapter.py    # Bridges tool_router to the harness contract
├── demo.py               # Toy 60-tool catalog demo of the POC
├── metrics.py            # Recall/precision/token-savings eval on the toy set
├── router-testkit/       # Hackathon input kit — DO NOT MODIFY
│   ├── catalog/          #   64 mock tools · 13 clusters · dup/version metadata
│   ├── test_cases/       #   29 scored cases across 11 categories
│   ├── registry_guardrail/#  Intake validator fixtures
│   └── harness/          #   run_benchmark.py (scorer) + baseline_router.py
├── bench/                # Additional benchmarks (baseline, mutations, guardrail)
├── specs/                # Spec-driven-development plan (SPEC-001..011)
├── docs/                 # PRDs and the hackathon problem statement
└── tests/                # Unit tests (populated as specs are executed)
```

## Baseline

| Router                                | Cases passed | Token savings |
|---------------------------------------|--------------|----------------|
| Testkit reference `baseline_router`   | 7 / 29 (24%) | 94% |
| `tool_router` (this repo, current)    | 5 / 29 (17%) | 90% |

The gap is deliberate and mapped in `specs/`. See
[`specs/README.md`](./specs/README.md) for the priority-ordered plan
(P0 → P1 → P2) with acceptance criteria pinned to specific `TC-*`
cases. Target after P0: ≥ 20/29. After P1: ≥ 24/29.

## Quick start

```bash
# Install (uv-managed)
uv sync

# Score the reference baseline
cd router-testkit/harness
python run_benchmark.py --verbose

# Score this project's router
python run_benchmark.py --router testkit_adapter:MyRouter --verbose
```

The harness has no third-party dependencies; standard-library only.
The router itself can run on `HashingEmbedder` (zero deps) or opt-in
`SentenceTransformerEmbedder` via `uv sync --extra semantic`.

## Where to read next

1. **`specs/README.md`** — the executable plan. Every spec is pickable
   independently and has acceptance criteria tied to testkit case IDs.
2. **`router-testkit/README.md`** — the grading contract (authoritative).
3. **`tool_router/README.md`** (below `tool_router/`) — the current POC
   library docs. Will be reframed in SPEC-011.
4. **`docs/hackathon-definition.md`** — the problem statement.
5. **`docs/PRD.md`** — the initial PRD.

## Toolchain

- Python ≥ 3.14 (pinned in `.python-version`, managed by uv)
- Ruff for lint + format
- Pytest for unit tests (see `pyproject.toml`'s `[dev]` extra)

## Dev

```bash
uv run ruff check .
uv run ruff format .
uv run pytest
```
