Metadata-Version: 2.5
Name: toolrank
Version: 0.2.0
Summary: Tool retrieval for LLM agents with hundreds of tools: an embedding model plus small learned heads picks the few tools a request needs, served over MCP and REST.
Project-URL: Homepage, https://github.com/yasinyaman/toolrank
Project-URL: Documentation, https://yaman.dev/toolrank/
Project-URL: Repository, https://github.com/yasinyaman/toolrank
Project-URL: Issues, https://github.com/yasinyaman/toolrank/issues
Project-URL: Changelog, https://github.com/yasinyaman/toolrank/blob/main/CHANGELOG.md
Author: Toolrank contributors
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: agents,embeddings,llm,mcp,retrieval,tool-retrieval,tool-search
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: bm25s>=0.2
Requires-Dist: numpy>=1.26
Provides-Extra: anthropic
Requires-Dist: anthropic>=1.9; extra == 'anthropic'
Provides-Extra: clm
Requires-Dist: torch>=2.2; extra == 'clm'
Provides-Extra: data
Requires-Dist: datasets>=2.19; extra == 'data'
Requires-Dist: huggingface-hub>=0.23; extra == 'data'
Provides-Extra: dev
Requires-Dist: anthropic>=1.9; extra == 'dev'
Requires-Dist: faiss-cpu>=1.15; extra == 'dev'
Requires-Dist: langchain-core>=1.0; extra == 'dev'
Requires-Dist: llama-index-core>=0.14; extra == 'dev'
Requires-Dist: mcp<3,>=2.2; extra == 'dev'
Requires-Dist: openai>=3.22; extra == 'dev'
Requires-Dist: pgvector>=0.4; extra == 'dev'
Requires-Dist: psycopg[binary]>=3.2; extra == 'dev'
Requires-Dist: pystemmer>=2.2; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: pyyaml>=6; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: faiss
Requires-Dist: faiss-cpu>=1.15; extra == 'faiss'
Provides-Extra: langgraph
Requires-Dist: langchain-core>=1.0; extra == 'langgraph'
Provides-Extra: llamaindex
Requires-Dist: llama-index-core>=0.14; extra == 'llamaindex'
Provides-Extra: lora
Requires-Dist: accelerate>=1.0; extra == 'lora'
Requires-Dist: peft>=0.14; extra == 'lora'
Requires-Dist: torch>=2.2; extra == 'lora'
Requires-Dist: transformers>=4.51; extra == 'lora'
Provides-Extra: mcp
Requires-Dist: mcp<3,>=2.2; extra == 'mcp'
Provides-Extra: openai
Requires-Dist: openai>=3.22; extra == 'openai'
Provides-Extra: openapi
Requires-Dist: pyyaml>=6; extra == 'openapi'
Provides-Extra: pgvector
Requires-Dist: pgvector>=0.4; extra == 'pgvector'
Requires-Dist: psycopg[binary]>=3.2; extra == 'pgvector'
Provides-Extra: stem
Requires-Dist: pystemmer>=2.2; extra == 'stem'
Description-Content-Type: text/markdown

# toolrank

[![PyPI](https://img.shields.io/pypi/v/toolrank.svg)](https://pypi.org/project/toolrank/)
[![CI](https://github.com/yasinyaman/toolrank/actions/workflows/ci.yml/badge.svg)](https://github.com/yasinyaman/toolrank/actions/workflows/ci.yml)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](https://github.com/yasinyaman/toolrank/blob/main/LICENSE)
[![Heads](https://img.shields.io/badge/heads-Hugging%20Face-yellow.svg)](https://huggingface.co/yasinyaman/toolrank-heads-qwen3-emb-8b)

**Tool retrieval for LLM agents with hundreds of tools.** Instead of putting every tool definition
into the prompt, toolrank picks the few a request needs, with an embedding model
(Qwen3-Embedding-8B, trained further on tool retrieval), and serves them to your agent over MCP or
REST. Every retriever it ships is measured on the same public benchmarks.

Status: **alpha (0.2)**; interfaces may still change. Documentation: <https://yaman.dev/toolrank/>

## Quick start

Index MCP servers or OpenAPI specs, then serve them to any MCP client as two tools, `search_tools`
and `call_tool`:

```bash
pip install "toolrank[mcp]"
toolrank ingest mcp --server time="uvx mcp-server-time" --out tools/
toolrank serve --data tools/                          # MCP at http://127.0.0.1:8765/mcp, REST at /v1
```

toolrank ranks with its backbone, Qwen3-Embedding-8B trained further on tool retrieval
(`yasinyaman/toolrank-emb-8b`), served by vLLM (`--emb-url`, by default
`http://127.0.0.1:8091/v1`). On a GPU host, Docker runs both:

```bash
cd deploy/docker && cp .env.example .env              # set TOOLRANK_API_KEY
docker compose run --rm toolrank ingest mcp --config /config/toolrank.json --out /data
docker compose up -d
```

The [quick start](https://yaman.dev/toolrank/quickstart/) connects Claude Code, Claude Desktop and REST clients.

## Results

Retrieval quality on three benchmarks, every number from `toolrank eval` under ToolRet's protocol
(the reports are in [`docs/results/`](https://github.com/yasinyaman/toolrank/tree/main/docs/results);
`scripts/readme_table.py` checks each one before it prints it).

<!-- results:start -->
| Retriever | ToolRet NDCG@10 | ToolRet NDCG@10 cat-macro | LiveMCPBench Recall@5 | MCP-Zero top-1 |
| --- | ---: | ---: | ---: | ---: |
| BM25, without instruction | 29.01 | 22.24 | 31.68 | 80.44 |
| BM25, with instruction | 39.27 | 36.41 | 22.92 | 45.63 |
| Qwen3-Embedding-8B | 51.11 | 46.54 | 50.82 | 78.19 |
| Qwen3-Embedding-8B + toolrank heads v0.1 | 54.03 | 47.13 | 53.03 | 79.87 |
| Qwen3-Embedding-8B in FP8 + toolrank heads v0.1 | 53.94 | 47.27 | **53.48** | 79.51 |
| toolrank backbone v0.2 (Qwen3-Embedding-8B + LoRA) | 58.90 | 54.36 | 52.06 | **88.57** |
| toolrank backbone v0.2 in FP8 (the default) | **59.02** | **54.53** | 52.06 | 87.71 |
| NV-Embed-v1 ([ToolRet paper](https://arxiv.org/abs/2503.01763)) | — | 42.71 | — | — |
| gte-Qwen2-1.5B-instruct ([ToolRet paper](https://arxiv.org/abs/2503.01763)) | — | 45.96 | — | — |
| StackOne v2, a fine-tuned 109M BGE-base ([StackOne](https://www.stackone.com/blog/autoresearch-charged-action-search/)) | — | 54.40 | — | — |

- ToolRet: 7,961 queries over 44,453 tools, top 100 over the whole corpus. *NDCG@10* is the micro-average of the paper's released code; *cat-macro* is the paper's own aggregation (the mean of the web, code and customized categories) and the only column with published numbers. Our BM25 reproduces the paper's BM25s within 0.1 (22.24 / 36.41 against 22.32 / 36.46).
- LiveMCPBench (94 queries, 525 tools) and MCP-Zero (2,792 tools): the tool text includes the MCP server's name (`toolrank data server-names`). One LiveMCPBench query is about one point. MCP-Zero ships no queries: ours were written by Qwen3-8B, one per tool (`toolrank data pull mcp-zero`), so its column does not compare with the MCP-Zero paper. Top-1 is Precision@1.
- *With instruction*, each query carries its task's instruction (ToolRet) or a generic one (the MCP sets), as the embedding model is served; the generic instruction costs BM25 on the MCP sets. BM25 is bm25s without stemming, the paper's setting.
- The heads (29.9M parameters, `docs/heads/MODEL_CARD.md`) were trained on ToolRet's training pairs, so ToolRet is in-domain for them and the MCP sets are not. On MCP-Zero, BM25 without instruction still wins at top-1: each generated query opens with a `server:` line that usually names the server, and exact matching rewards that.
- The toolrank backbone v0.2 (`docs/backbone/MODEL_CARD.md`) is Qwen3-Embedding-8B with a LoRA trained on 20,000 of ToolRet's training pairs, served without heads (the v0.1 heads cost it 1–2 points). Its checkpoint was picked on MCP-Zero, so that column is its selection set; ToolRet is in-domain, LiveMCPBench is held out. Beyond these sets, on generated requests over a GitHub + Stripe catalogue it gains 12–20 NDCG@10 points on tasks that need two or three tools and ties with the base model on requests for one tool (the model card has the numbers).
- FP8: the bf16 weights quantized as vLLM loads them (`--quantization fp8`). Every column is within a point of bf16, at about half the weight memory and batch-1 latency.
- Reproduce: `bash scripts/readme_results.sh` where the backbone is served (reports in `docs/results/`; `EMB_URL`, `EMB_MODEL`, `TAG` and `ROWS` select another endpoint and rows), then `uv run python scripts/readme_table.py --write`.
<!-- results:end -->

## What's inside

- **Ingestion** of MCP servers (stdio and streamable HTTP) and OpenAPI 3.x specs; a re-run syncs only
  what changed. [Guide](https://yaman.dev/toolrank/guides/ingest/)
- **Search and serve**: adaptive K, a persistent vector index (numpy, FAISS HNSW or pgvector), an MCP
  proxy with two tools, a REST API, API keys and a usage log. [Guide](https://yaman.dev/toolrank/guides/serve/)
- **Agent platforms**: toolrank as Claude's (`tool_reference`) and OpenAI's (client-side
  `tool_search`) tool search. [Guide](https://yaman.dev/toolrank/guides/platforms/)
- **Frameworks**: LangGraph (langgraph-bigtool), LlamaIndex agents and the LiteLLM proxy.
  [Guide](https://yaman.dev/toolrank/guides/frameworks/)
- **Fine-tuning**: heads trained on your own request-to-tool pairs, the epoch picked on a dev set.
  [Guide](https://yaman.dev/toolrank/guides/finetune/)
- **Benchmarks**: ToolRet, LiveMCPBench and MCP-Zero with BM25, dense and head scorers (and CLM, for
  comparison). [Benchmarks](https://yaman.dev/toolrank/benchmarks/)
- **Docker**: the `toolrank` image for amd64 and arm64, compose files with vLLM, and a Dockerfile
  that puts vLLM and toolrank in one container.
  [Guide](https://yaman.dev/toolrank/guides/docker/)

## Why

- **Tool definitions are expensive context.** Anthropic measured 58 tools at about 55K tokens per
  request and reports tool-selection accuracy falling past 30 to 50 tools. RAG-MCP lifted
  selection accuracy from 13.6% to 43.1% on a large MCP set by retrieving tools first.
- **Hosted tool searches are tied to one model provider or cloud, and most are lexical.** toolrank
  is model-agnostic and runs on your own hardware.
- **Small heads on a frozen embedding model** (29.9M parameters for the request and the tool side
  together). You embed your tools once and rank with one dot product. The heads start as the
  identity, so heads fine-tuned on your data start from the base model's quality, not below it.

## Feedback

Tried it? [Tell us how it went](https://github.com/yasinyaman/toolrank/issues/new?template=feedback.yml):
what you set up, what worked and what did not. Questions and ideas go to
[Discussions](https://github.com/yasinyaman/toolrank/discussions).

## Contributing

Set-up, tests and conventions are in
[`CONTRIBUTING.md`](https://github.com/yasinyaman/toolrank/blob/main/CONTRIBUTING.md). The weekly
reports behind every number, in Turkish, are in
[`docs/reports/`](https://github.com/yasinyaman/toolrank/tree/main/docs/reports).

## License

Apache-2.0 ([`LICENSE`](https://github.com/yasinyaman/toolrank/blob/main/LICENSE)). [`NOTICE`](https://github.com/yasinyaman/toolrank/blob/main/NOTICE) credits what toolrank takes from CLM (the head
architecture), ToolRet (its task metadata) and MCP-Zero (its query prompts);
[`THIRD_PARTY_NOTICES.md`](https://github.com/yasinyaman/toolrank/blob/main/THIRD_PARTY_NOTICES.md) lists the dependencies, models and container
images. Contributions are welcome: see [`CONTRIBUTING.md`](https://github.com/yasinyaman/toolrank/blob/main/CONTRIBUTING.md) and the
[code of conduct](https://github.com/yasinyaman/toolrank/blob/main/CODE_OF_CONDUCT.md); report vulnerabilities as [`SECURITY.md`](https://github.com/yasinyaman/toolrank/blob/main/SECURITY.md) says.
