Metadata-Version: 2.4
Name: bit-jev
Version: 0.13.12
Summary: A Jev/Kev-style decision model on Microsoft's 1.58-bit BitNet b1.58 backbone
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/Zeaulo/bit-jev
Project-URL: Documentation, https://github.com/Zeaulo/bit-jev#快速开始
Project-URL: Model, https://huggingface.co/jinghao1632/bit-jev-2b-distilled
Project-URL: ModelScope, https://www.modelscope.cn/models/JingHao9616/bit-jev-2b-distilled
Requires-Python: <3.13,>=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: transformers<5,>=4.48
Requires-Dist: huggingface_hub
Requires-Dist: modelscope-hub<1,>=0.4.3
Requires-Dist: gradio<6,>=5
Provides-Extra: train
Requires-Dist: torch; extra == "train"
Requires-Dist: peft==0.14.0; extra == "train"
Requires-Dist: accelerate>=1.3; extra == "train"
Requires-Dist: numpy<2; extra == "train"
Requires-Dist: safetensors>=0.5; extra == "train"
Requires-Dist: sentencepiece; extra == "train"
Provides-Extra: serve
Requires-Dist: fastapi; extra == "serve"
Requires-Dist: uvicorn; extra == "serve"
Requires-Dist: torch; extra == "serve"
Requires-Dist: peft==0.14.0; extra == "serve"
Requires-Dist: accelerate>=1.3; extra == "serve"
Requires-Dist: numpy<2; extra == "serve"
Requires-Dist: safetensors>=0.5; extra == "serve"
Provides-Extra: data
Requires-Dist: datasets; extra == "data"
Requires-Dist: torch; extra == "data"
Requires-Dist: numpy<2; extra == "data"
Provides-Extra: benchmark
Requires-Dist: psutil>=5.9; extra == "benchmark"
Dynamic: license-file

# bit-jev

**Install, open a local decision page, then build your own structured requests.** bit-jev scores explicit options over a BitNet backbone and returns answers and probabilities without generating answer text token by token.

![bit-jev ternary inputs and decision engine](https://raw.githubusercontent.com/Zeaulo/bit-jev/main/docs/figures/decision-engine.png)

[GitHub documentation](https://github.com/Zeaulo/bit-jev) · [Hugging Face model](https://huggingface.co/jinghao1632/bit-jev-2b-distilled) · [ModelScope model](https://www.modelscope.cn/models/JingHao9616/bit-jev-2b-distilled) · [中文说明](https://github.com/Zeaulo/bit-jev/blob/main/README.md)

## Try it

```bash
pip install bit-jev -i https://pypi.org/simple --upgrade
python -m bit_jev.demo
```

The second command opens a local Gradio page with Choice, Noul, and Score tabs. Choice and Score let you add or remove items within a two-to-four-item range. The first submitted question downloads the 1.19 GB GGUF from Hugging Face or, if that connection fails, ModelScope; later requests reuse the loaded model. Installation and page startup do not download weights. Use `python -m bit_jev.demo --source modelscope` to select ModelScope directly, or `--device gpu` for Vulkan. Use `--once` for the former one-shot JSON command.

The Windows x64 wheel includes precompiled CPU and Vulkan GPU runners for AVX2 processors. Inference on those machines needs no Git, CMake, C++ compiler, or Vulkan SDK. Vulkan still needs a compatible graphics driver and its `vulkan-1.dll` runtime. Other platforms build the native runner on demand and require [Git](https://git-scm.com/install/), [CMake 3.28+](https://cmake.org/download/), and a C++17 compiler ([Windows C++ Build Tools](https://learn.microsoft.com/cpp/build/vscpp-step-0-installation)). Source builds of Vulkan also need its SDK; CUDA builds need a CUDA Toolkit. Both precompiled programs use pinned BitNet and llama.cpp source with the ReLU² runtime patch and carry their MIT license notices.

## Resident inference

```python
from bit_jev.gguf import BitJev

request = {
    "state": "A customer reports a duplicate charge.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this?",
            "criteria": {"billing": "Payment and refund issues", "shipping": "Delivery issues"},
        }
    },
}

with BitJev.from_pretrained(device="cpu", threads=8) as model:
    result = model.infer(request)
    print(result["answers"], result["latency_ms"])
```

Use `device="gpu"` for Vulkan or `device="cuda"` for an NVIDIA CUDA build. GPU requests fail clearly if a backend or visible GPU is unavailable. Pass `source="modelscope"` to select ModelScope directly, or pass a local model directory for offline inference. `binary="/path/to/bit-jev-cpu"` selects an existing native runner. The model stays loaded for repeated `infer()` calls; `latency_ms` reports native compute only, excluding download, build, loading, encoding, and IPC.

The CLI accepts UTF-8 JSONL input:

```bash
bit-jev --device cpu --input requests.jsonl --output results.jsonl
```

Training and distillation dependencies are optional: `pip install 'bit-jev[train]'`. The native runner evaluates one causal row per question; multiple questions repeat the shared state. The repository's Apache-2.0 license covers source code, while the checkpoint has no standalone open-weights license. The model card documents Yelp training-data provenance and its unresolved permission request.
