Metadata-Version: 2.4
Name: litellm-cost
Version: 0.1.0
Summary: Standalone LLM cost estimation: model pricing data + token cost math, extracted from LiteLLM with zero litellm dependency.
Author-email: AKN Runtime <akn-runtime@users.noreply.github.com>
License: MIT
Project-URL: Homepage, https://github.com/akn-runtime/litellm-cost
Project-URL: Documentation, https://github.com/akn-runtime/litellm-cost#readme
Project-URL: Source, https://github.com/akn-runtime/litellm-cost
Project-URL: Upstream, https://github.com/BerriAI/litellm
Keywords: llm,cost,token,pricing,estimation,litellm
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries
Classifier: Topic :: System :: Systems Administration
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: tiktoken
Requires-Dist: tiktoken>=0.7; extra == "tiktoken"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-cov>=5.0; extra == "dev"
Dynamic: license-file

# litellm-cost

**Standalone LLM cost estimation: model pricing data + token cost math, with zero `litellm` dependency.**

[![PyPI - Version](https://img.shields.io/pypi/v/litellm-cost.svg)](https://pypi.org/project/litellm-cost/)
[![PyPI - Python Version](https://img.shields.io/pypi/pyversions/litellm-cost.svg)](https://pypi.org/project/litellm-cost/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Tests](https://img.shields.io/badge/tests-114%20passing-brightgreen.svg)]()

---

## Why this package exists

[LiteLLM](https://github.com/BerriAI/litellm) is a 100+ provider LLM gateway — but its cost
estimation logic is trapped inside a 9,700-line god-module (`litellm/utils.py`) and paid for
with a heavy dependency bill (pydantic, httpx, openai, jinja2, …). Two design choices make it
painful to reuse in isolation:

1. **Import-time network fetch** — `import litellm` pulls the pricing JSON from GitHub by
   default, which fails or blocks in offline / air-gapped environments and adds a supply-chain
   surface.
2. **Dependency bloat** — the cost path only needs a data dictionary and arithmetic, yet
   installing `litellm` drags in the full gateway stack.

`litellm-cost` extracts exactly the reusable core — the pricing data dictionary and the
token-to-cost math — into a **zero-dependency** package that:

- never touches the network on import (or at all, unless you explicitly ask it to refresh),
- exposes a small, typed, documented API,
- ships an embedded pricing snapshot with its own `data_version` for traceability,
- keeps the upstream O(1) case-insensitive model lookup and CJK-aware token heuristics.

## Installation

```bash
pip install litellm-cost
```

Optional extras:

```bash
pip install "litellm-cost[tiktoken]"   # tiktoken accelerator for token estimation
pip install "litellm-cost[dev]"        # pytest + pytest-cov for development
```

Python **3.9 – 3.12** supported. No third-party dependencies in the core package.

## Quick start

### Estimate the cost of a call

```python
from litellm_cost import get_cost

# gpt-4o: $2.5e-06 / input token, $1e-05 / output token
cost = get_cost("gpt-4o", input_tokens=1250, output_tokens=350)
print(f"${cost:.6f}")   # $0.006625

# provider-prefixed model keys work as in LiteLLM
cost = get_cost("groq/llama-3.1-8b-instant", input_tokens=1000, output_tokens=500)

# cached tokens are billed at the model's cache-read price
# (defaults to 50% of the input price when the data has no explicit value)
cost = get_cost("gpt-4o", input_tokens=1000, cache_read_tokens=900)

# or pass an OpenAI-style usage block directly
cost = get_cost("gpt-4o", usage={"prompt_tokens": 100, "completion_tokens": 50})
```

### Estimate tokens (no tokenizer dependency)

```python
from litellm_cost import estimate_tokens

estimate_tokens("Hello, world!")                        # heuristic, CJK-aware
estimate_tokens("你好世界，这是一个测试。")                # CJK text handled correctly
estimate_tokens([{"role": "user", "content": "Hi"}])    # OpenAI-style messages

# opt-in tiktoken accelerator (silently falls back to the heuristic
# when tiktoken is not installed)
estimate_tokens("Hello, world!", use_tiktoken=True, model="gpt-4o")
```

### Query pricing data

```python
from litellm_cost import get_pricing, list_providers, data_version

info = get_pricing("gpt-4o")
print(info["input_cost_per_token"])    # 2.5e-06
print(info["max_input_tokens"])        # 128000

providers = list_providers()           # full 100+ provider roster
print(data_version())                  # e.g. "2025.06.17.1"
```

### Handle errors explicitly

```python
from litellm_cost import LitellmCostError, UnknownModel, MissingPricing, ContextWindowExceededError

try:
    get_cost("not-a-real-model", input_tokens=10)
except UnknownModel as exc:
    print(exc)   # unknown model: 'not-a-real-model'. It is not present in the embedded pricing data...

try:
    get_cost("gpt-4o", input_tokens=10**9, check_context_window=True)
except ContextWindowExceededError:
    pass  # opt-in: off by default, cost estimation should not fail on hypothetical usage
```

### Keep pricing data fresh

The embedded snapshot is versioned and traceable. To update it from LiteLLM's upstream
`model_prices_and_context_window.json` (stdlib only, no third-party dependencies):

```bash
python -m litellm_cost.refresh --check          # dry run: report without writing
python -m litellm_cost.refresh                  # fetch, merge, bump data_version
```

Or programmatically:

```python
from litellm_cost.refresh import refresh_pricing_data

summary = refresh_pricing_data()
print(summary)   # {"data_version": ..., "models": ..., "providers": ..., "written": True}
```

The refresh pipeline inherits LiteLLM's own supply-chain guards: the fetched payload must be a
non-empty JSON object, and the refresh is always an explicit, opt-in action — never an import
side effect.

## API overview

| Function | Purpose |
|---|---|
| `get_cost(model, ...)` | Compute the USD cost of a call from token counts |
| `estimate_tokens(source, ...)` | Estimate tokens for text or OpenAI-style messages |
| `get_pricing(model)` | Query the pricing-info dict for a model (defensive copy) |
| `list_providers()` | Sorted list of supported providers (100+) |
| `cost_per_token(model)` | `(input_cost_per_token, output_cost_per_token)` tuple |
| `data_version()` | Version string of the embedded pricing data |

Full signatures, parameter tables, error semantics and examples:
**[docs/API.md](docs/API.md)**.

## Relationship to LiteLLM

This package is a **decomposition** of the cost-estimation logic found in LiteLLM
(`litellm.cost_calculator`, `litellm.utils`, `litellm_core_utils.token_counter`, and the
`model_prices_and_context_window.json` data dictionary), stripped of runtime dependencies
(network I/O, provider routing, logging, telemetry). The extraction boundary review covers
LiteLLM **v1.85.1**; the package is versioned independently of upstream.

What is *not* copied: the router, the proxy server, provider request adapters, the
logging/telemetry stack, streaming processors, and the import-time network fetch.

## License

MIT — see [LICENSE](LICENSE). The LICENSE retains the upstream attribution notice: the pricing
data schema and token-to-cost math are derived from LiteLLM (MIT), and the embedded pricing
JSON is upstream community-maintained data with its source URL recorded in the bundle.

## Development

```bash
git clone https://github.com/akn-runtime/litellm-cost
cd litellm-cost
pip install -e ".[dev]"
pytest
```

To rebuild the distribution artifacts:

```bash
pip install build
python -m build          # produces sdist + wheel in dist/
```
