Metadata-Version: 2.4
Name: theclassifier
Version: 0.1.0
Summary: Natural language classification on CPU and GPU, with domain adapters and LoRA fine-tuning
Author: TrueState
License-Expression: Apache-2.0
Project-URL: Homepage, https://theclassifier.ai
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch<3,>=2.6
Requires-Dist: transformers<6,>=4.56
Requires-Dist: peft<1,>=0.17
Requires-Dist: accelerate<2,>=1
Requires-Dist: huggingface-hub<2,>=0.34
Requires-Dist: safetensors>=0.5
Provides-Extra: serve
Requires-Dist: fastapi<1,>=0.115; extra == "serve"
Requires-Dist: uvicorn<1,>=0.30; extra == "serve"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: httpx; extra == "dev"
Requires-Dist: build; extra == "dev"
Requires-Dist: twine; extra == "dev"
Dynamic: license-file

# The Classifier

A general-purpose natural language classifier with domain-specific LoRA adapters.
Run locally on CPU or NVIDIA GPU, serve the same `/v1/classify` request format as
the hosted service, or fine-tune an adapter on your own labelled examples.

## Release status

This is the initial 0.1.0 package. The configured model repository is
`TrueState/theclassifier`. The checkpoints are staged privately; authenticate with
Hugging Face and obtain repository access before using named-version downloads.
A local checkpoint also works.
Model weights are not included in the Python package. This code is Apache-2.0;
checkpoint and dataset licences are separate.

## Install

From this checkout, until the PyPI release is published:

```sh
pip install './python[serve]'
```

After release:

```sh
pip install 'theclassifier[serve]'
```

Python 3.10+ is required. Install a PyTorch build suitable for your platform first
if needed; CUDA support depends on that build and your NVIDIA driver. CPU uses
float32; CUDA uses bfloat16 when supported, otherwise float32. No FlashAttention
extension is required. `auto` selects CUDA when available and CPU otherwise.
Apple MPS is not currently supported.

## Classify locally

```python
from theclassifier import Classifier

classifier = Classifier()  # general; downloads once, then uses the local cache
result = classifier.classify(
    'Please close my account.',
    choices=['billing', 'cancellation', 'none'],
    question='What is the customer asking for?',
)
print(result['results'][0]['selected'])
```

Use `Classifier(checkpoint='./checkpoints/general', device='cpu')` to load an
existing local checkpoint. A checkpoint contains `adapter/`, `tokenizer/` and
`head.safetensors` (or the existing `head.pt` format). The pretrained backbone is
resolved from the adapter configuration. A local adapter can still require a
backbone download if the backbone is not already cached.

The full question format supports multiple decisions and optional criteria:

```python
result = classifier.classify(
    'Please close my account.',
    questions=[{
        'question': 'What is the customer asking for?',
        'choices': [
            {'choice': 'cancellation', 'criteria': 'Requests to close an account.'},
            {'choice': 'none', 'criteria': 'None of the other choices apply.'},
        ],
    }],
    version='intent',
)
```

Versions: `general`, `intent`, `sentiment`, `triage`, `cause`, `codify`,
`complaints`, `stars`. Selecting a new version checks the Hugging Face cache,
downloads missing files, then loads the adapter and its scoring head onto the
shared backbone. Selection and inference are serialized to keep the adapter and
head paired. A failed download raises an error; it never silently falls back to
another model.

```python
classifier.load_adapter('custom', checkpoint='./checkpoints/my-adapter')
result = classifier.classify('example', choices=['yes', 'no'], version='custom')
```

Use `model_repo=`, `revision=` and `cache_dir=` or the
`THECLASSIFIER_MODEL_REPO` environment variable to override download settings.
Pin `revision` to a commit for reproducibility. `offline=True` disables network
loading, including the base model. Pre-download adapters with:

```sh
theclassifier download --version general
theclassifier download --version intent
```

These commands download adapters and tokenizers; initialize the classifier once
online to cache the backbone as well. The repository layout is
`<version>/adapter/`, `<version>/tokenizer/`, `<version>/head.safetensors`.

Default maximum input length is 2,048 tokens per candidate; longer inputs retain
the first token and tail, matching the existing prompt format. Set `max_length`
explicitly for longer records. `batch_size` controls inference candidate chunks.
This initial portable runtime uses direct candidate scoring, not shared-prefix
KV caching. It has not been benchmarked against the optimized production server.

## Serve on CPU or GPU

```sh
theclassifier serve --checkpoint ./checkpoints/general --device cpu
# On a CUDA machine:
API_KEY=your-serving-key theclassifier serve --device cuda --host 0.0.0.0
```

Default binding is `127.0.0.1:8090`. Binding outside localhost requires `API_KEY`.
Use one process per GPU; run behind a TLS proxy for remote clients. This standalone
server provides inference and authentication, not the hosted service's billing,
quotas or rate limiting. Set resource limits in your deployment.

```sh
curl http://127.0.0.1:8090/v1/classify \
  -H 'Content-Type: application/json' \
  -d '{"state":"Close my account","questions":[{"question":"Intent?","choices":[{"choice":"cancel"},{"choice":"none"}]}]}'
```

Include `Authorization: Bearer YOUR_KEY` when `API_KEY` is configured.
Responses contain `results` (with `question`, `selected`, `detailed_results`),
`version`, `input_tokens` and `latency_ms`.

## Fine-tune with LoRA

Each JSONL line is one complete choice menu, with an exact matching answer label:

```json
{"state":"Close my account","question":"Intent?","choices":[{"choice":"cancel","criteria":"Requests to close an account"},{"choice":"none"}],"answer":"cancel"}
```

```sh
# Continue the general checkpoint to create a specialist:
theclassifier finetune --version general --data train.jsonl \
  --output checkpoints/my-adapter --device cuda

# Train an adapter and scoring head from the base model:
theclassifier finetune --base-model Qwen/Qwen3-1.7B-Base \
  --data train.jsonl --output checkpoints/from-base --rank 16 --alpha 16
```

Or in Python:

```python
from theclassifier import finetune

report = finetune(
    'train.jsonl', 'checkpoints/my-adapter',
    checkpoint='checkpoints/general', device='auto', epochs=1,
    learning_rate=2e-4, grad_accum=16,
)
```

Training freezes the backbone, learns LoRA updates to `q_proj` and `v_proj`, and
trains a float32 scalar head using cross-entropy across each menu. Whole menus
are processed together, gradients accumulate between menus, and gradient
checkpointing reduces activation memory. Large choice menus may still require
smaller input lengths or more memory. CPU training is supported but slow.

New adapters default to rank 16, alpha 16 and dropout 0.05. Continuing an existing
adapter preserves its LoRA configuration. A new AdamW optimizer is created with
a constant learning rate; this is fine-tuning, not exact training-state resume.
The optimizer configuration differs from the original research runs. Output
folders must be empty. Saved files include the adapter, head, tokenizer,
`classifier.json` and `training.json`. Keep evaluation data separate and evaluate
before using a new checkpoint; the training loss is not a quality benchmark.

## Development and release

```sh
cd python
pip install -e '.[serve,dev]'
pytest
python -m build
python -m twine check dist/*
# With PyPI credentials configured locally:
python -m twine upload dist/*
```

The tests build a tiny random Qwen model locally: no production weights, private
data or external inference service is needed. All six tests passed on the NVIDIA
workstation, including CUDA inference. Real v1.1 checkpoints were also smoke-tested
on CUDA, and the general model was checked on CPU against the original scoring
implementation. These checks establish execution compatibility, not accuracy.
Publishing the package does not upload model weights. See `RELEASE.md` for the
checkpoint and publication steps.
