Metadata-Version: 2.5
Name: qev
Version: 0.2.1
Summary: LAYA-inspired typed decisions over text and images with Qwen3.5-2B
Project-URL: Homepage, https://github.com/ken-jo/qev
Project-URL: Documentation, https://github.com/ken-jo/qev#readme
Project-URL: Repository, https://github.com/ken-jo/qev
Project-URL: Issues, https://github.com/ken-jo/qev/issues
Project-URL: Model, https://huggingface.co/ken-jo/qev
Project-URL: Dataset, https://huggingface.co/datasets/ken-jo/qev-data
Author: ken-jo
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: classification,multimodal,probabilities,typed-decisions
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: <3.13,>=3.12
Requires-Dist: accelerate==1.15.0
Requires-Dist: fastapi==0.141.1
Requires-Dist: filelock<5,>=3.16
Requires-Dist: gradio==6.29.0
Requires-Dist: huggingface-hub<2,>=1.0
Requires-Dist: numpy==2.5.3
Requires-Dist: pillow==12.3.0
Requires-Dist: pydantic==2.13.5
Requires-Dist: safetensors==0.8.0
Requires-Dist: torch==2.10.0
Requires-Dist: torchvision==0.25.0
Requires-Dist: transformers==5.17.0
Requires-Dist: uvicorn==0.54.0
Description-Content-Type: text/markdown

# QEV

[![Multilingual](https://img.shields.io/badge/languages-Multilingual-2563eb)](#language-support)

[![Sponsor on GitHub](https://img.shields.io/badge/Sponsor-GitHub-ea4aaa?logo=githubsponsors&logoColor=white)](https://github.com/sponsors/ken-jo)
[![PyPI](https://img.shields.io/pypi/v/qev)](https://pypi.org/project/qev/)

**Your evidence. Your criteria. A decision with probabilities.**

[Model on Hugging Face](https://huggingface.co/ken-jo/qev) ·
[Dataset](https://huggingface.co/datasets/ken-jo/qev-data)

QEV is an open multimodal decision model **inspired by
[LAYA](https://huggingface.co/convaiinnovations/laya)** and built with the text and vision
backbone of **[Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B)**. Give it text, a photo,
or both, then describe the decision you need. It returns candidate probabilities, a typed
answer and an abstention signal in one batched backbone forward, with zero generated
answer tokens.

LAYA's request-defined typed decisions motivated the interface. QEV adds visual evidence
through Qwen's existing multimodal backbone, learned language adapters, option readouts
and a condition-modulated binding head. The Qwen vision encoder is frozen. LAYA weights
are not embedded in this checkpoint, and the current training recipe is supervised
adaptation with calibration. RLCD remains a research direction. QEV is independently
maintained; upstream model attribution is preserved.

## Language support

**Multilingual inputs** use Qwen3.5-2B's multilingual text backbone. The playground and
documentation are in English. Published QEV task and calibration evaluations focus on
English; equivalent accuracy across languages has not been established.

## Install and run

The **QEV 0.2.1 Python SDK includes the English playground**, six licensed example
photographs, image resolution controls, and the inference API. Use Python 3.12.

### pip

```sh
python -m pip install qev==0.2.1
qev playground
```

Open **http://127.0.0.1:7860**. The first launch downloads the pinned QEV adaptation and
Qwen3.5-2B backbone (about 4.6 GB), then loads the model. Later launches reuse the cache.
CUDA is used when available; `--device cpu` and `--device cuda` select a device explicitly.
Published on [PyPI](https://pypi.org/project/qev/0.2.1/) through GitHub Trusted Publishing.
The same runtime is also available as a wheel with the Hugging Face model.

### uv

Run the published package without cloning the repository:

```sh
uvx --python 3.12 --from qev==0.2.1 qev playground
```

Or use the source checkout and its frozen CUDA dependency lock:

```sh
git clone https://github.com/ken-jo/qev.git
cd qev
uv sync --frozen
uv run qev playground
```

For NVIDIA CUDA 12.8 wheels with `uvx`, add
`--torch-backend cu128` before `--from`. With pip, install the
matching CUDA-enabled PyTorch and torchvision first if the default installation is CPU-only.

### Python SDK

```python
from qev import DecisionRequest, load

model = load()  # Downloads on first use; reuses the model cache afterwards.
request = DecisionRequest.model_validate({
    "state": {"text": "I was charged twice. Please refund the duplicate payment."},
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {"billing": "Payments and refunds", "technical": "Software faults"},
        }
    },
})
print(model.predict(request)["answers"])
```

For JSON files: `qev predict --request examples/request.json`. For an HTTP API:
`qev serve --image-root examples`. Both prepare the model on first use. `qev download`
can fetch weights in advance; `--offline` requires a complete cache. Set `QEV_HOME` for
QEV's persistent data directory or `QEV_CACHE_DIR` for the Hugging Face cache.

The Python distribution contains code, UI and example photos. Model tensors are downloaded
separately. The inference weights and 38 internal `veyra` modules remain the evaluated
QEV 0.1.1 model; SDK 0.2.1 adds installation and application features.
See [the playground guide](https://github.com/ken-jo/qev/blob/main/docs/PLAYGROUND.md)
and [API schema](https://github.com/ken-jo/qev/blob/main/docs/API.md).

## Decisions you define at request time

Route a support message, assess an ordered severity level, or ask whether a photo meets
a written condition. Change the candidate descriptions to change the task. The model
scores the choices supplied with the request; its output slots do not represent a fixed
catalog of classes.

These are intended uses for domain evaluation. Published results below establish the
current scope, including failures in visual reasoning and unfamiliar tasks.

## What you can ask

| Type | Request-specific definition | Output |
| --- | --- | --- |
| `choice` | 2-16 candidate descriptions | Candidate probabilities and selected label |
| `score` | 2-16 ordered level descriptions | Level probabilities and expected level |
| `noul` | A proposition, with true/false criteria | True/false probabilities |

Accepts text, one image, or both; 1-4 questions per request. Candidate meanings are supplied
at inference time. This does not imply reliable generalization to every unseen task.

```mermaid
flowchart LR
    E[Text and optional image] --> Q[Qwen3.5-2B with language LoRA]
    C[Question and candidate descriptions] --> Q
    Q --> R[Option readout and condition binding]
    R --> P[Type-specific calibration]
    P --> A[Probabilities, typed answer, abstention]
```

## QEV and LAYA on the same inputs

| English test | LAYA English | LAYA Typed Decisions | QEV 0.1.1 |
| --- | ---: | ---: | ---: |
| Typed decisions, 2,000 questions | 36.05% | 76.95% | 77.00% |
| News topic, 400 examples | 95.00% | 95.25% | 82.50% |
| Emotion, 400 examples | 58.75% | 60.00% | 50.25% |

![Measured accuracy with 95% intervals](https://raw.githubusercontent.com/ken-jo/qev/main/reports/release-comparison/accuracy.png)

Typed-decisions is an adapted benchmark for QEV and the specialist. Their one-question
accuracy difference does not establish an advantage (paired 95% interval: -1.85 to +1.95
percentage points). News and emotion are **zero-shot relative to QEV's audited adaptation
data**; unknown backbone pretraining overlap remains possible. LAYA reports news in its
training mix and emotion held out. All models received the same inputs, without truncation.

LAYA is smaller and faster in this comparison. On typed-decisions, its specialist has
lower probability errors and 23.71 ms resident p50 versus QEV's 77.70 ms on the same GPU.
QEV adds image input; that capability has its own evaluations and limitations.
See [the full protocol, probability metrics and results](https://github.com/ken-jo/qev/blob/main/docs/LAYA_COMPARISON.md).

## A photo, three typed answers

The release includes a reproducible photograph example: one plastic bottle, six material
candidates, a three-level handling policy, and a true/false proposition.

<img src="https://raw.githubusercontent.com/ken-jo/qev/main/examples/photograph/item.jpg" alt="TrashNet verification photograph of a plastic bottle" width="360" />

| Type | Recorded output |
| --- | --- |
| `choice` | Plastic: **0.9361** probability |
| `score` | Expected policy level: **0.4254** on a 0-2 scale |
| `noul` | Glass, paper or plastic: **0.7903** probability true |

One batched forward, zero generated answer tokens. These are actual outputs on a previously
inspected verification fixture, not an independent accuracy estimate. Photo: TrashNet,
Gary Thung, MIT. [Request, output, image and reproduction](https://github.com/ken-jo/qev/tree/main/examples/photograph).

## Image and workflow evaluation

| Evaluation | Result | What it covers |
| --- | ---: | --- |
| Fresh procedural workflows | 69.38% | 1,440 questions, 480 groups, 3 authored families |
| Fresh CIFAR-10 photograph guard | 95.83% | 600 low-resolution image questions |
| Fresh SNLI | 87.78% | 450 questions |
| Fresh BANKING77 | 85.50% | 462 sampled eight-candidate questions, not standard 77-way classification |
| Official typed-decisions regression | 77.00% | 2,000 questions; 30.80% coverage under abstention |
| Resident local photo HTTP p95 | 114.94 ms | RTX 4060 Ti 8 GB; serial, one photo/question, six candidates |

The timing excludes loading, WAN transport and concurrency. It is not a general latency SLA.
On the shifted final population, accepted uncertain requests had **45.57% expected error**;
overall ECE was **12.58%**. Calibration does not guarantee correctness.

**Visual reasoning remains limited:** 0 of 72 exploratory 2048 games reached 2048.
A narrow balanced board-cell diagnostic scored 5/28 for images and 21/28 for text.
These failures are published alongside the successful measurements. Read the
[model card](https://github.com/ken-jo/qev/blob/main/MODEL_CARD.md) and [evaluation report](https://github.com/ken-jo/qev/blob/main/docs/EVALUATION.md) before using it.

## What you download

| Artifact | Hosted on | Purpose |
| --- | --- | --- |
| QEV adaptation and decision heads | Hugging Face `ken-jo/qev` | The learned QEV weights and calibration |
| Qwen3.5-2B backbone | Hugging Face `Qwen/Qwen3.5-2B` | The pinned upstream text and vision model |
| `qev` Python SDK | PyPI and release wheel | Loading, typed inference, downloads, serving and playground |
| Research and application source | GitHub `ken-jo/qev` | Training scripts, evaluations and local interfaces |

Installing the SDK installs code, the playground, samples and dependencies. First use
fetches the two weight components; `qev download` can prepare them in advance. The SDK's package size is not the model's size.

## Size and precision

The inference backbone has **2.213B parameters**. The checkpoint adds **7.992M stored
adapter/readout parameters** in a **32.01 MB safetensors file**. GPU inference uses a
BF16 backbone and FP32 decision readouts; the CPU path uses FP32. LoRA is merged into
the backbone for inference. Upstream weights download separately (about **4.55 GB**).
See [exact counts and tensor types](https://github.com/ken-jo/qev/blob/main/docs/MODEL_SIZE.md).

## Playground and API

`qev playground` starts the packaged English interface. It includes text/image presets,
six sample photographs, choice/score/noul, image-resolution choices, candidate probabilities,
abstention and the exact request/response JSON. It runs locally; no public Space is created.

```sh
qev playground --host 0.0.0.0 --port 7860
qev serve --image-root examples --host 127.0.0.1 --port 8000
```

The first command makes the UI accessible through your PC's IP on a trusted network.
The second serves `POST /v1/systemone`; image paths resolve under the selected image root.
The playground serializes model requests. Send API requests one at a time.
These are local interfaces; they do not provide a public multi-tenant service.

The supported playground is the English interface included in the package. The Windows
helpers in `apps/playground/` launch the same interface. The retired 2048 experiments remain
research records and are not part of the playground.

## Agent skill

The [QEV skill](skills/qev/SKILL.md) lets Codex and other skill-capable agents compose
text/image judgments, call QEV, preserve probability and abstention fields, and choose
an authorized next step. It includes a validated helper, three-type image example and
runtime instructions. A resident model server avoids loading weights for every decision.

The skill is distributed separately from the pip runtime, in `skills/qev` and the release's
`qev-skill-0.2.1.zip`. Copy that complete folder into your agent's skills directory and use
`$qev`. See [installation and examples](docs/AGENT_SKILL.md). The skill adds orchestration;
it does not change the model or establish Jev-equivalent performance.

## Data, evidence and attribution

- [Training and reproduction](https://github.com/ken-jo/qev/blob/main/docs/TRAINING.md): staged training, provenance and limits.
- [Dataset publication](https://github.com/ken-jo/qev/blob/main/docs/DATA.md): 15 historical corpus configurations, original splits,
  image hashes and source-specific licenses. Configurations overlap; do not concatenate them.
- [Research evidence](https://github.com/ken-jo/qev/blob/main/reports/README.md): successful and failed experiments.
- [Release history](https://github.com/ken-jo/qev/blob/main/CHANGELOG.md) and [future research](https://github.com/ken-jo/qev/blob/main/ROADMAP.md).

Code and adaptation weights: Apache-2.0. Qwen retains its upstream Apache-2.0 attribution.
Dataset licenses differ by source. LAYA and JEV inspired typed decision interfaces; their
weights and code are not included, and no affiliation or equivalent performance is claimed.
See [NOTICE](https://github.com/ken-jo/qev/blob/main/NOTICE) and [source licenses](https://github.com/ken-jo/qev/blob/main/docs/data-licenses/).

## Support

If QEV is useful, [star the repository](https://github.com/ken-jo/qev) or visit
[GitHub Sponsors](https://github.com/sponsors/ken-jo). `qev star --open` opens
the repository for you to choose; using QEV does not require a star or donation.

[GitHub: ken-jo/qev](https://github.com/ken-jo/qev)
