Metadata-Version: 2.5
Name: jev-checker
Version: 0.2.0
Summary: A semantic Python code checker powered by TypeSafe Jev
Requires-Python: >=3.11
Requires-Dist: httpx<1,>=0.27
Requires-Dist: pathspec<2,>=1
Requires-Dist: python-dotenv<2,>=1
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: pyrefly>=1.3.2; extra == 'dev'
Requires-Dist: pytest-cov>=5; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.9; extra == 'dev'
Requires-Dist: twine>=6; extra == 'dev'
Description-Content-Type: text/markdown

# jev-checker

A Python semantic checker that adds Jev powered review diagnostics alongside Ruff and Pyrefly.
Install from PyPI:

```bash
python -m pip install jev-checker
```

## Quick start

Jev loads a `.env` file from the current directory or a parent directory. Put one
provider key in that file; shell environment variables take precedence:

```bash
# .env
TYPESAFE_API_KEY=...
# Or use OPENROUTER_API_KEY=... and pass --provider openrouter

# Run with TypeSafe
jev check .

# Run with OpenRouter
jev check . --provider openrouter
```

The `jev-checker check` executable is an alias for `jev check`.

Check committed changes against the merge base of a branch:

```bash
jev check . --diff origin/main
```

`--diff` reviews commits only. It does not include uncommitted working tree changes.
The source excerpts selected for review are sent to the chosen provider after
credential pattern redaction. Redaction cannot identify every secret; review your
provider's data policy before using the checker on sensitive code.

## Diagnostics

The default output resembles other checkers:

```text
src/auth.py:42:5: error: a protected operation may lack authorization [JEV201]

1 error, 0 warnings, 0 infos
```

`error` fails the check; warnings and infos do not unless `--warnings-as-errors`
is set. Configure rule levels in `pyproject.toml`:

```toml
[tool.jev]
provider = "typesafe"
select = ["JEV1", "JEV2", "JEV3", "JEV4", "JEV5", "JEV6", "JEV7"]

[tool.jev.levels]
JEV305 = "warning"
JEV601 = "error"
JEV704 = "info"
JEV503 = "off"
```

The checker keeps diagnostic level, Jev probability, confidence, and impact score
as separate values. Use `--output-format full` for evidence details, `--show` to
print every screening score and follow-up decision, or `json` for a machine
readable report including screening scores. Supported formats are `concise`,
`full`, `json`, and `github`. Exit codes are 0 for no errors, 1 for blocking diagnostics, and 2 when
the requested analysis cannot be completed. `--exit-zero` affects only code 1.
`--fail-on-unresolved` also returns 2 for candidates without a reliable conclusion;
`--exit-zero` does not override that explicit policy.
JSON reports also distinguish supported, dismissed, context-limited, and incomplete
candidates while omitting raw source excerpts.

## Built-in rules

- `JEV1` correctness, `JEV2` security, `JEV3` reliability, `JEV4` compatibility
- `JEV5` test evidence, `JEV6` performance, `JEV7` observability

Security, correctness, reliability, and compatibility rules default to errors.
Test, performance, and observability rules default to warnings. Use `--select`,
`--ignore`, and `--rule-level RULE=LEVEL` to adjust the run.

## Reading `--show`

Each `path:start-end` heading is a source region. The score beside each rule is
the model's screening probability from 0 to 1. `*` means the score reached
`screen-threshold` (default `0.70`) and was queued for follow-up; it is a signal,
not a confirmed finding or a warning by itself. A diagnostic is emitted only
after follow-up finds sufficiently confident, reachable evidence. These scores
are model judgments, not calibrated guarantees.

### Rule IDs

| ID | What it screens for |
| --- | --- |
| JEV101 | A condition or branch handles the wrong cases. |
| JEV102 | A value is sent to the wrong field, argument, or recipient. |
| JEV103 | A state update breaks an established invariant. |
| JEV104 | An accepted boundary input produces an incorrect result. |
| JEV105 | An operation uses a dependent result before it is ready. |
| JEV201 | A protected operation lacks required authorization. |
| JEV202 | A user can access another user's or tenant's data or operations. |
| JEV203 | Untrusted input reaches an interpreter or execution mechanism unsafely. |
| JEV204 | Untrusted input escapes a filesystem or network boundary. |
| JEV205 | Sensitive data reaches an unauthorized output or recipient. |
| JEV206 | A default configuration leaves a required safeguard disabled. |
| JEV301 | A failure path leaks an acquired resource. |
| JEV302 | Concurrent work can update shared state inconsistently. |
| JEV303 | A wait or resource acquisition can stall progress indefinitely. |
| JEV304 | Cancellation leaves work or external effects inconsistent. |
| JEV305 | Retrying after failure can duplicate a non-idempotent effect. |
| JEV401 | A public contract conflicts with supplied consumers. |
| JEV402 | A persisted or exchanged format conflicts with supplied readers. |
| JEV403 | Exception or exit-code behavior conflicts with supplied consumers. |
| JEV404 | Code needs a capability missing from a declared Python version or platform. |
| JEV501 | Important behavior lacks targeted related-test evidence. |
| JEV502 | An important failure or recovery path lacks test evidence. |
| JEV503 | An authorization boundary lacks an access-denial test. |
| JEV504 | An important boundary input lacks targeted test evidence. |
| JEV505 | A test assertion may pass even when the relevant behavior is wrong. |
| JEV601 | A blocking operation may run on an asynchronous event loop. |
| JEV602 | Repeated external calls may have an equivalent batch alternative. |
| JEV603 | Memory use may grow without a bound on input size. |
| JEV604 | An invariant expensive calculation may be repeated unnecessarily. |
| JEV701 | An important failure may end without an observable signal. |
| JEV702 | A failed operation may be reported as successful. |
| JEV703 | Distributed work may lose an identifier needed to connect its events. |
| JEV704 | A diagnostic may contradict the operation's actual result. |

The prefixes group rules: `JEV1` correctness, `JEV2` security, `JEV3`
reliability, `JEV4` compatibility, `JEV5` test evidence, `JEV6` performance,
and `JEV7` observability. Their default levels are shown under [Built-in rules](#built-in-rules).

### Profiles and follow-up decisions

Source reviews screen each named function independently, including methods,
async functions, and nested functions. Executable code outside functions is also
screened. Every focused request includes the complete Python file as context,
including its original line breaks. Localization selects evidence inside that
focus; verification also receives the complete file. Candidates for the same rule
in different functions are investigated independently. `--show` displays the
qualified function names alongside screening scores and follow-up decisions.
Related tests are selected by module imports or exact
test-file naming conventions and are supplied in full. Source profiles also receive
the full file. Credential literals remain redacted while preserving their positions.

The checker imposes no byte or line cap on this context. The provider enforces its
own model limits: Jev currently documents 32k tokens for state plus the longest
question and 64k tokens for the entire request ([model limits](https://docs.typesafe.ai/models)).
Bytes are not tokens. Rejected requests produce an incomplete review and exit code
2; source is never silently truncated to make a request fit. External files and
contracts are only available when explicitly supplied through the review inputs.

`File profiles` show the file category and its confidence, then review priority
and that score's confidence. Review priority runs from 0 (routine) to 3
(specialist or immediate review); it is not a diagnostic severity.

`Follow-up decisions` list each rule selected for closer review. `p=` repeats
its screening probability. `supported (diagnostic_emitted)` means a diagnostic
was produced; `dismissed` means follow-up did not confirm an issue;
`insufficient_context` means confidence or evidence was too low; and
`incomplete` means the follow-up could not finish. The reason in parentheses
gives more detail, for example `low_location_confidence` means Jev could not
pin the concern to a source region confidently enough to emit a diagnostic.

## Checker pipeline

Run each tool as a separate step so its result stays visible:

```bash
ruff check .
ruff format --check .
pyrefly check
jev check . --diff origin/main --output-format json > jev-report.json
```

Provide the diff base in the CI checkout. Keep provider credentials in the CI
secret store and run authenticated reviews only in trusted jobs. `jev check .
--dry-run` reports file and region counts without using credentials or the network.
There is no default cap on API attempts, evidence follow-ups, or run duration;
all selected units and candidates are reviewed, with at most three concurrent
requests. Use `--max-requests N`, `--max-followups N`, or `--max-duration SECONDS`
to impose an explicit cap. A cap or deadline that interrupts review produces an
incomplete result, not a clean check.
The local SQLite cache fingerprints the complete
file content. Unchanged files reuse judgments; any content edit automatically
rechecks all review stages for that file, including functions whose own bodies
did not change, because their contracts and callers may have changed. Other
unchanged files keep their cache. Changes to the diff's
before version or a supplied related test also invalidate the affected file's
review. A timestamp change alone does not invalidate identical content.

The cache expires after 24 hours by default, including version-family aliases
such as `typesafe/jev-1.13`. Explicit
`latest`, `auto`, and `preview` aliases bypass it. For reproducible evaluations,
use a provider-supported dated model identity and record the returned model.
Normal use requires no `--no-cache`; that flag only forces a fresh run of unchanged
content. The provider controls request pricing and rate limits.

## Development

```bash
python -m pip install -e '.[dev]'
ruff check .
ruff format --check .
pyrefly check
pytest
pytest --cov=jev_review --cov-report=term-missing
```

The tests use simulated providers and do not make paid API calls. See
[`docs/judgment-catalog.md`](docs/judgment-catalog.md) for rule prompts and evidence
criteria, and [`docs/implementation-plan.md`](docs/implementation-plan.md) for the
implementation checklist and remaining release work.


## Unresolved reviews and evidence integrity

The default output now exposes suspicious code that could not be confirmed:

```text
No confirmed diagnostics; review has unresolved or unfinished work.
0 errors, 0 warnings, 0 infos
0 remote requests, 2 cache hits; models: jev-fixture
Unresolved: 1 candidate(s); this is not a clean review.
  sample.py [JEV101]: low_location_confidence
```

Use the following command to make unresolved work fail a verification pipeline:

```bash
jev check src/ --provider openrouter --fail-on-unresolved
```

The equivalent configuration is `fail-on-unresolved = true` under `[tool.jev]`.
Without that policy, a completed run with only unresolved candidates retains exit
0 for compatibility. Confirmed errors return 1; operational failures, exhausted
budgets, and the explicit unresolved policy return 2.

Screening remains at 0.70, verification at 0.80, and location/mechanism confidence
at 0.70 by default. Intermediate verification probabilities remain unresolved;
a sufficiently negative judgment dismisses the concern. Impact confidence describes
uncertainty about severity and is reported separately from defect confirmation.
Optional reviewer routing failures preserve confirmed diagnostics.

Evidence is split into bounded spans with exact ranges and content hashes. Oversized
or omitted evidence is recorded; localization no longer silently sends only a prefix.
Containing definitions and called local contracts are collected without executing
reviewed code. Broad source locations are marked as evidence anchors in full output.

### JSON schema migration

Reports now use `schema_version: 2`. Existing diagnostic, screening, candidate,
usage, and operational `status` fields remain available. Consumers must accept
version 2 and inspect `review_status` and `unresolved_count`: operational completion
alone does not mean every candidate was resolved. New `traces` expose judgments,
thresholds, evidence metadata, and remote/cache provenance without raw source or
credentials. Cached judgments are reprocessed with current policy; prompt/state
changes invalidate their cache keys. `token_usage_status` distinguishes reported
usage from unavailable or partial usage and runs with no remote requests; token totals
sum only known remote usage.

## Semantic evaluation

Ordinary pytest checks use offline providers. The opt-in suite contains 31 buggy/fixed
pairs, separate tuning/held-out splits, and static-analysis comparison controls.
It measures actual model detection separately from pipeline tests:

```bash
python -m tools.evaluate_semantics --dry-run --split all
python -m tools.evaluate_semantics --provider openrouter --split tuning \
  --case success_inversion --case negative_transfer --case missing_authorization \
  --case failed_replacement --repetitions 3 --max-total-requests 100
```

Live evaluations send the repository-owned fixtures to the selected provider and
may incur charges. The request budget applies to the entire run, including retries.
They run without cache unless `--cache-replay` is explicit. Unresolved bugs count
as missed detections; unattempted cases remain visible. A targeted rule evaluation
is a tuning aid and cannot satisfy the end-to-end release gates.

See [evaluation instructions](docs/semantic-evaluation.md) and the
[reliability implementation checklist](docs/semantic-review-reliability-plan.md)
for measured limitations and remaining release criteria.

### Experimental direct review

From the repository checkout, evaluate a direct per-function runtime judgment:

```bash
uv run python -m tools.scan_direct_review benchmarks/realistic_review \
  --provider openrouter --output .jev/direct-review.json
uv run python -m tools.evaluate_semantics --strategy direct --split all \
  --repetitions 3 --max-total-requests 1000 --output .jev/direct-corpus.json
```

This experiment uses complete-file context and reports suggestions at probability
0.70 without a second defect-confirmation gate. Uncertain positions are function
anchors. It always uses fresh judgments and is separate from `jev check`.
Reserved recall was 38/60 by expected rule, or 44/60 after verifying alternative
rule classifications, with 3/60 corrected runs receiving false alerts. Three
large-module findings were reproduced locally. The production quality gate
remains unmet; see [the experiment report](docs/direct-review-experiment.md).

### Manual semantic scenarios

[The scenario corpus](benchmarks/semantic_review/README.md) adds 24 buggy/fixed
pairs with obvious and subtle defects, varying impact, and executable local
witnesses. It covers permissions, tenant isolation, money, concurrency, retries,
formats, performance, and observability. Scan an individual example or directory:

```bash
uv run jev-checker check benchmarks/semantic_review/buggy --provider openrouter --show
```

The catalog explains each expected defect and its corresponding corrected file.
Use the evaluation runner's `--corpus benchmarks/semantic_review --split all`
option for repeated measurements with neutral paths and negative controls.

### Pipeline benchmark

[`benchmarks/pipeline_review`](benchmarks/pipeline_review/README.md) adds 5
buggy/fixed module pairs forming one small document-processing system, each
defect in a different rule family (JEV301, JEV402, JEV604, JEV702, JEV205).
Offline witnesses run with ordinary `pytest`; live detection runs through the
same `tools.evaluate_semantics --corpus benchmarks/pipeline_review` entry point.

### Calibration benchmark

[`benchmarks/calibration_review`](benchmarks/calibration_review/README.md)
targets implicit (undocumented) authorization requirements, the pattern that
motivated lowering `verification_threshold`'s default to `0.80`. Use
`--cache-replay` with a different `--verification-threshold` to reprocess
cached judgments at a new threshold without new API calls.
