Metadata-Version: 2.4
Name: zanii-llm-client
Version: 0.1.0
Summary: Client for the Zanii LLM API: OpenAI-compatible inference with a verifiable receipt for every call
Author-email: Zanii <info@zanii.agency>
License-Expression: MIT
Project-URL: Homepage, https://llm.zanii.agency
Project-URL: Documentation, https://llm.zanii.agency/docs
Project-URL: Source, https://github.com/vigilancetrent/zanii-llm
Project-URL: Issues, https://llm.zanii.agency/contact
Keywords: llm,openai,inference,receipts,audit,uae
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.12
Classifier: Intended Audience :: Developers
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: verify
Requires-Dist: zanii>=0.24.0; extra == "verify"
Provides-Extra: dev
Requires-Dist: ruff; extra == "dev"
Dynamic: license-file

# zanii-llm-client

Python client for [Zanii LLM](https://llm.zanii.agency): OpenAI-compatible inference where every
call can produce a signed receipt on a public ledger.

```bash
pip install zanii-llm-client            # add [verify] to check proofs client-side
```

```python
from zanii_llm_client import Zanii

z = Zanii()                                       # reads ZANII_LLM_API_KEY
answer = z.chat("Summarise this claim in two sentences.", model="glm-4.7-flash")

print(answer.text)
print(answer.cost_aed, "AED")
print(z.verify(answer.receipt).ok)
```

No dependencies. A dependency in a client package becomes a dependency in every application that
installs it.

## You may not need this

The API speaks the OpenAI protocol, so the `openai` package works with one changed base URL and will
keep working:

```python
from openai import OpenAI
client = OpenAI(api_key="zlk_live_...", base_url="https://llm.zanii.agency/v1")
```

This package adds the parts no OpenAI client knows about.

| | |
|---|---|
| `answer.receipt`, `z.verify(...)` | the proof for a call, checked against the public ledger |
| `InsufficientCredit` | a typed 402 carrying the top-up link, not an opaque error |
| `z.usage()`, `z.balance()` | reconcile our invoice against your own records |
| `max_spend_micro` | a client-side ceiling; the request is never made |
| `thinking=False` | skip the model's reasoning when the question does not need it |

## Reasoning costs money

These models reason before they answer. Reasoning tokens are billed like any other and come out of
`max_tokens`. Measured on `glm-4.7-flash`, "name one thing Sharjah is known for" costs 412 output
tokens and 4.4 seconds with reasoning on, and 22 tokens and 0.6 seconds with it off, for the same
answer.

```python
z.chat("Name one thing Sharjah is known for.", model="glm-4.7-flash", thinking=False)
```

If `answer.text` comes back empty, `answer.empty_because` says why and `answer.reasoning_tokens`
says where the budget went.

## Streaming

```python
for piece in z.stream("Write three lines about Dubai.", model="glm-4.7-flash"):
    print(piece, end="", flush=True)

print(z.last)                                      # the receipt, once the stream ends
```

## Verifying a receipt

```python
check = z.verify(answer.receipt)
assert check.ok
assert check.detail["verified_locally"]            # True with the [verify] extra installed
```

Verification fetches a Merkle inclusion proof from `ledger.zanii.agency` and checks it against a
signed tree head. It does not call Zanii LLM at all.

## Parity

The TypeScript client [`@zanii/llm`](https://www.npmjs.com/package/@zanii/llm) mirrors this one
method for method. A conformance suite runs the same cases through both and fails on a difference.

---

© Zanii, United Arab Emirates · [llm.zanii.agency](https://llm.zanii.agency) · info@zanii.agency
