Metadata-Version: 2.4
Name: nat-csoai-evidence
Version: 0.1.0
Summary: NeMo Agent Toolkit evaluator: verify a CSOAI signed evidence batch offline and report each event's state (CONSISTENT / DIVERGENT / UNMEASURED ...), never a grade.
Author: CSOAI Ltd
License: Apache-2.0
Project-URL: Homepage, https://councilof.ai/connect/
Project-URL: Signing key (DID), https://csoai.org/.well-known/did.json
Keywords: nemo-agent-toolkit,nat,evaluator,evidence,ed25519,opentelemetry
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Security
Requires-Python: <3.14,>=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: nvidia-nat-core<2,>=1.9
Requires-Dist: nvidia-nat-eval<2,>=1.9
Requires-Dist: cryptography>=41
Provides-Extra: eval
Requires-Dist: nvidia-nat-eval[full]<2,>=1.9; extra == "eval"
Requires-Dist: langchain-core>=0.3; extra == "eval"
Dynamic: license-file

# nat-csoai-evidence: NeMo Agent Toolkit evaluator for signed evidence

Apache-2.0. This is a `nat.components` plugin that registers the evaluator `_type: csoai_evidence`. For each dataset item it does two things:

1. It verifies a CSOAI signed evidence bundle **offline** against a pinned DID document. The bundle is made of `batch.json`, `batch.signed.json`, `events.jsonl` and one `event_id`.
2. It reports that event's state word: CONSISTENT, DIVERGENT, PARTIAL, UNMEASURED, UNCHECKABLE or NOT_DISCRIMINATING.

If any byte of the bundle was changed, the result is INVALID. If the signing key is not pinned, the result is UNVERIFIABLE_KEY.

## Install

```sh
pip install "nat-csoai-evidence[eval]"                    # nvidia-nat-core + nvidia-nat-eval[full] >= 1.9 (the [eval] extra brings what the nat eval command needs, incl. langchain-core, which nvidia-nat-eval 1.9.0 imports but does not declare)
curl -o did.json https://csoai.org/.well-known/did.json    # pin the issuer's keys
```

In your NAT eval config:

```yaml
eval:
  evaluators:
    evidence:
      _type: csoai_evidence
      did_json: ./did.json
```

```sh
nat eval --config_file eval.yml --skip_workflow
```

Each dataset item's generated answer is a signed-evidence bundle. The source distribution carries a five-item fixture set (`fixtures/`), signed with a **published test key** so it attests nothing. Each item returns its expected result: CONSISTENT, DIVERGENT, UNMEASURED, INVALID (a one-word edit to the events), UNVERIFIABLE_KEY (a key that is not pinned).

**No score.** Every item's `score` is `None`, so `average_score` is `None`. The evaluator's `reasoning` carries the OTel `gen_ai.evaluation.result` attributes, and `score.value` is absent unless a number was actually measured.

**Errors are not zeros.** NAT's `BaseEvaluator` records an evaluator error as a score of 0.0. This plugin does not subclass it. An error here is reported as UNCHECKABLE with no score.

**Limit.** VALID shows who signed these bytes. It does not show that any claim inside is true.

This plugin is our code. NVIDIA holds the toolkit; this is not an NVIDIA integration and no NVIDIA team has reviewed it. More: https://councilof.ai/connect/
