Metadata-Version: 2.4
Name: gemmadecision
Version: 0.1.0
Summary: Typed local decisions for PydanticAI, with Rust HTTP serving and batched Gemma inference
Project-URL: Homepage, https://github.com/rjn32s/gemmadecision
Project-URL: Repository, https://github.com/rjn32s/gemmadecision
Project-URL: Issues, https://github.com/rjn32s/gemmadecision/issues
Project-URL: Model, https://huggingface.co/rajan2k/GemmaDecision-270M
Author: Rajan Shukla
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: decision-model,gemma,granian,pydantic-ai,vllm
Classifier: Development Status :: 4 - Beta
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: fastapi<0.137,>=0.133
Requires-Dist: granian<3,>=2.8
Requires-Dist: httpx<1,>=0.28
Requires-Dist: huggingface-hub<2,>=1.33
Requires-Dist: numpy<3,>=2
Requires-Dist: pydantic-ai-slim<3,>=2.51
Requires-Dist: pydantic<3,>=2.12
Requires-Dist: safetensors<1,>=0.8
Requires-Dist: torch<2.15,>=2.13
Requires-Dist: transformers<6,>=5.17
Provides-Extra: http
Requires-Dist: fastapi<0.137,>=0.133; extra == 'http'
Requires-Dist: granian<3,>=2.8; extra == 'http'
Provides-Extra: pydantic-ai
Requires-Dist: pydantic-ai-slim<3,>=2.51; extra == 'pydantic-ai'
Provides-Extra: serve
Requires-Dist: fastapi<0.137,>=0.133; extra == 'serve'
Requires-Dist: granian<3,>=2.8; extra == 'serve'
Requires-Dist: huggingface-hub<2,>=1.33; extra == 'serve'
Requires-Dist: numpy<3,>=2; extra == 'serve'
Requires-Dist: safetensors<1,>=0.8; extra == 'serve'
Requires-Dist: torch<2.15,>=2.13; extra == 'serve'
Requires-Dist: transformers<6,>=5.17; extra == 'serve'
Provides-Extra: test
Requires-Dist: asgi-lifespan<3,>=2; extra == 'test'
Requires-Dist: build<2,>=1; extra == 'test'
Requires-Dist: fastapi<0.137,>=0.133; extra == 'test'
Requires-Dist: granian<3,>=2.8; extra == 'test'
Requires-Dist: pytest-asyncio<2,>=1; extra == 'test'
Requires-Dist: pytest<10,>=8; extra == 'test'
Requires-Dist: twine<7,>=6; extra == 'test'
Provides-Extra: vllm
Requires-Dist: fastapi<0.137,>=0.133; extra == 'vllm'
Requires-Dist: granian<3,>=2.8; extra == 'vllm'
Requires-Dist: transformers<6,>=5.17; extra == 'vllm'
Requires-Dist: vllm==0.30.0; extra == 'vllm'
Description-Content-Type: text/markdown

# gemmadecision

Small, local decisions in one Python call. Powered by
[GemmaDecision-270M](https://huggingface.co/rajan2k/GemmaDecision-270M).

## Install

```bash
pip install gemmadecision
```

Python 3.11+. No API key or server needed. The first call downloads the model
(about 0.5 GB); later calls reuse it. CPU, NVIDIA CUDA and Apple Silicon are
selected automatically. After downloading, inference runs locally.

## Use it in your code

```python
from gemmadecision import decide

team = decide("I was charged twice", choices=["billing", "technical"])
print(team)  # selected choice, as a string
```

For more specific decisions, give each choice a description:

```python
team = decide(
    "The same payment appears twice on my statement.",
    choices={
        "billing": "Handle charges, payments and refunds",
        "technical": "Handle crashes, login errors and app problems",
    },
    question="Which support team should handle this request?",
)
```

Need the full ranking? Use `rank(...)` with the same arguments. It returns
ordered candidates, scores and derived probabilities.

## PydanticAI

```python
from typing import Literal
from pydantic_ai import Agent
from gemmadecision import GemmaDecisionModel

agent = Agent(
    GemmaDecisionModel.local(),
    output_type=Literal["billing", "technical"],
)
result = agent.run_sync("I was charged twice")
print(result.output)
```

This uses PydanticAI's native decision-model interface. Boolean, enum, rubric
and finite Pydantic fields are supported.
[More PydanticAI examples](https://github.com/rjn32s/gemmadecision/blob/main/docs/pydantic_ai.md).

## Serve it

```bash
gemmadecision serve
```

That starts the Rust-powered Granian HTTP server at `http://127.0.0.1:8700`.
Interactive API documentation is available at `/docs`.

```python
from gemmadecision import DecisionClient

with DecisionClient() as client:
    result = client.decide(
        "I was charged twice",
        candidates={"billing": "Payments and refunds", "technical": "App problems"},
    )
    print(result.choice)
```

`AsyncDecisionClient` works with `await`. To connect PydanticAI to the server,
use `GemmaDecisionModel()` instead of `.local()`.

## More control when you need it

- [Hardware settings, offline use, HTTP API and batching](https://github.com/rjn32s/gemmadecision/blob/main/docs/advanced.md)
- [vLLM backend for Linux/CUDA](https://github.com/rjn32s/gemmadecision/blob/main/docs/vllm.md)
- [Performance measurements and reproduction](https://github.com/rjn32s/gemmadecision/blob/main/docs/performance.md)
- [Docker and releases](https://github.com/rjn32s/gemmadecision/blob/main/docs/releasing.md)

The default uses batched PyTorch for model computation and Rust for HTTP
serving. vLLM is optional; no Rust compiler is needed to install the package.

This model chooses among supplied options; it does not generate open-ended
text. Its scores are rankings, and the derived probabilities are not a guarantee
of correctness. Maximums are 2,048 state/question tokens, 768 tokens per choice
and 64 choices. See the [model card](https://huggingface.co/rajan2k/GemmaDecision-270M)
for evaluation and limitations.

Code: Apache-2.0. Model weights: separate Gemma terms. See [NOTICE](https://github.com/rjn32s/gemmadecision/blob/main/NOTICE).
