Metadata-Version: 2.4
Name: bit-jev
Version: 0.8.9
Summary: A Jev/Kev-style decision model on Microsoft's 1.58-bit BitNet b1.58 backbone
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/Zeaulo/bit-jev
Project-URL: Documentation, https://github.com/Zeaulo/bit-jev#快速开始
Project-URL: Model, https://huggingface.co/jinghao1632/bit-jev-2b-distilled
Requires-Python: <3.13,>=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: transformers<5,>=4.48
Requires-Dist: huggingface_hub
Provides-Extra: train
Requires-Dist: torch; extra == "train"
Requires-Dist: peft==0.14.0; extra == "train"
Requires-Dist: accelerate>=1.3; extra == "train"
Requires-Dist: numpy<2; extra == "train"
Requires-Dist: safetensors>=0.5; extra == "train"
Requires-Dist: sentencepiece; extra == "train"
Provides-Extra: serve
Requires-Dist: fastapi; extra == "serve"
Requires-Dist: uvicorn; extra == "serve"
Requires-Dist: torch; extra == "serve"
Requires-Dist: peft==0.14.0; extra == "serve"
Requires-Dist: accelerate>=1.3; extra == "serve"
Requires-Dist: numpy<2; extra == "serve"
Requires-Dist: safetensors>=0.5; extra == "serve"
Provides-Extra: data
Requires-Dist: datasets; extra == "data"
Requires-Dist: torch; extra == "data"
Requires-Dist: numpy<2; extra == "data"
Provides-Extra: benchmark
Requires-Dist: psutil>=5.9; extra == "benchmark"
Dynamic: license-file

# bit-jev

bit-jev scores explicit options over a 1.58-bit BitNet backbone. Its I2_S GGUF inference path loads a resident model and returns structured answers, logits, and probabilities without generating answer tokens.

[GitHub documentation](https://github.com/Zeaulo/bit-jev) · [GGUF model and model card](https://huggingface.co/jinghao1632/bit-jev-2b-distilled) · [中文说明](https://github.com/Zeaulo/bit-jev/blob/main/README.md)

## Install

```bash
pip install bit-jev
```

The wheel contains Python code and native build sources. The 1.19 GB GGUF is downloaded from Hugging Face on first use. Native compilation requires [Git](https://git-scm.com/install/), [CMake 3.28+](https://cmake.org/download/), and a C++17 compiler ([Windows C++ Build Tools](https://learn.microsoft.com/cpp/build/vscpp-step-0-installation)). Git retrieves and verifies pinned BitNet and llama.cpp source and applies the ReLU² patch; normal inference does not need Git. Version 0.8.8 checks build tools before downloading the model. Vulkan GPU mode additionally requires a Vulkan SDK; CUDA mode requires a CUDA Toolkit. Neither model download nor compilation runs during `pip install`.

## Resident inference

```python
from bit_jev.gguf import BitJev

request = {
    "state": "A customer reports a duplicate charge.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "Payment and refund issues", "shipping": "Delivery issues"},
        }
    },
}

with BitJev.from_pretrained(device="cpu", threads=8) as model:
    result = model.infer(request)
    print(result["answers"], result["latency_ms"])
```

Use `device="gpu"` for Vulkan or `device="cuda"` for an NVIDIA CUDA build. GPU requests fail clearly if a backend or visible GPU is unavailable. A local model directory can replace the default Hugging Face repo, and `binary="/path/to/bit-jev-cpu"` can select an existing native runner. The model stays loaded for repeated `infer()` calls; `latency_ms` reports native compute only, excluding download, build, loading, encoding, and IPC.

The CLI accepts UTF-8 JSONL input:

```bash
bit-jev --device cpu --input requests.jsonl --output results.jsonl
```

Training and distillation dependencies are optional: `pip install 'bit-jev[train]'`. The native runner evaluates one causal row per question; multiple questions repeat the shared state. The repository's Apache-2.0 license covers source code, while the checkpoint has no standalone open-weights license. The model card documents Yelp training-data provenance and its unresolved permission request.
