Metadata-Version: 2.4
Name: styxx
Version: 7.48.0
Summary: Cognitive observability for LLM agents. @styxx.profile returns a per-step cognometric readout: drift, confabulation, refusal, sycophancy, phase transition, low trust, incoherence. Calibrated AUC: 0.998 hallucination (HaluEval-QA), 0.976 refusal (XSTest-GPT-4), 0.943 tool-call drift (BFCL v3) — register instruments with documented construct ceilings. Reference implementation of the Cognometric Fingerprint Spec v1.0. No torch, no GPU, no LLM in the loop; base install is numpy + scikit-learn.
Author-email: flobi <heyzoos123@gmail.com>
Maintainer-email: flobi <heyzoos123@gmail.com>
License: MIT
Project-URL: Homepage, https://styxx-org.netlify.app
Project-URL: Documentation, https://styxx-org.netlify.app
Project-URL: Fathom Lab, https://fathomlab-io.netlify.app
Project-URL: Telegram, https://t.me/STYXX_COMM
Project-URL: Research Paper, https://doi.org/10.5281/zenodo.19326174
Project-URL: @fathom_lab, https://x.com/fathom_lab
Project-URL: Source, https://github.com/fathom-lab/styxx
Project-URL: Issue Tracker, https://github.com/fathom-lab/styxx/issues
Keywords: llm,agent,agent-observability,cognitive,cognitive-profiler,cognometric-fingerprint,refusal-detection,tool-call-drift,sycophancy-detection,interpretability,ai-safety,langsmith,datadog,fathom,spec-v1.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24
Requires-Dist: scikit-learn>=1.3
Provides-Extra: sklearn
Requires-Dist: scikit-learn>=1.3; extra == "sklearn"
Provides-Extra: plate
Requires-Dist: matplotlib>=3.7; extra == "plate"
Requires-Dist: scipy>=1.10; extra == "plate"
Provides-Extra: openai
Requires-Dist: openai>=1.0; extra == "openai"
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.20; extra == "anthropic"
Provides-Extra: langchain
Requires-Dist: langchain-core>=0.1; extra == "langchain"
Provides-Extra: crewai
Requires-Dist: crewai>=0.1; extra == "crewai"
Requires-Dist: langchain-core>=0.1; extra == "crewai"
Provides-Extra: autogen
Requires-Dist: pyautogen>=0.2; extra == "autogen"
Provides-Extra: langsmith
Requires-Dist: langsmith>=0.1; extra == "langsmith"
Requires-Dist: langchain-core>=0.1; extra == "langsmith"
Provides-Extra: langfuse
Requires-Dist: langfuse>=2.0; extra == "langfuse"
Requires-Dist: langchain-core>=0.1; extra == "langfuse"
Provides-Extra: agent-card
Requires-Dist: Pillow>=10.0; extra == "agent-card"
Requires-Dist: matplotlib>=3.7; extra == "agent-card"
Provides-Extra: sense
Requires-Dist: psutil>=5.9; extra == "sense"
Provides-Extra: tier1
Requires-Dist: torch>=2.1; extra == "tier1"
Requires-Dist: transformers>=4.40; extra == "tier1"
Requires-Dist: transformer-lens>=2.0; extra == "tier1"
Provides-Extra: tier2
Requires-Dist: torch>=2.1; extra == "tier2"
Requires-Dist: transformers>=4.40; extra == "tier2"
Requires-Dist: transformer-lens>=2.0; extra == "tier2"
Requires-Dist: circuit-tracer>=0.4; extra == "tier2"
Provides-Extra: nli
Requires-Dist: torch>=2.0; extra == "nli"
Requires-Dist: transformers>=4.35; extra == "nli"
Requires-Dist: sentence-transformers>=2.2; extra == "nli"
Provides-Extra: mcp
Requires-Dist: mcp>=1.0.0; python_version >= "3.10" and extra == "mcp"
Provides-Extra: hf
Requires-Dist: transformers>=4.40; extra == "hf"
Requires-Dist: torch>=2.6; extra == "hf"
Provides-Extra: coherence
Requires-Dist: scipy>=1.10; extra == "coherence"
Provides-Extra: signing
Requires-Dist: cryptography>=41.0; extra == "signing"
Provides-Extra: guardrails
Requires-Dist: guardrails-ai>=0.4; extra == "guardrails"
Provides-Extra: llamaindex
Requires-Dist: llama-index-core>=0.10; extra == "llamaindex"
Provides-Extra: schema
Requires-Dist: jsonschema>=4.0; extra == "schema"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: ruff>=0.1; extra == "dev"
Requires-Dist: jsonschema>=4.0; extra == "dev"
Provides-Extra: test
Requires-Dist: pytest>=7.0; extra == "test"
Requires-Dist: ruff>=0.1; extra == "test"
Requires-Dist: jsonschema>=4.0; extra == "test"
Requires-Dist: mcp>=1.0.0; python_version >= "3.10" and extra == "test"
Requires-Dist: anthropic>=0.20; extra == "test"
Requires-Dist: openai>=1.0; extra == "test"
Requires-Dist: Pillow>=10.0; extra == "test"
Requires-Dist: matplotlib>=3.7; extra == "test"
Requires-Dist: langchain-core>=0.1; extra == "test"
Requires-Dist: psutil>=5.9; extra == "test"
Dynamic: license-file

<div align="center">

```
   ███████╗████████╗██╗   ██╗██╗  ██╗██╗  ██╗
   ██╔════╝╚══██╔══╝╚██╗ ██╔╝╚██╗██╔╝╚██╗██╔╝
   ███████╗   ██║    ╚████╔╝  ╚███╔╝  ╚███╔╝
   ╚════██║   ██║     ╚██╔╝   ██╔██╗  ██╔██╗
   ███████║   ██║      ██║   ██╔╝ ██╗██╔╝ ██╗
   ╚══════╝   ╚═╝      ╚═╝   ╚═╝  ╚═╝╚═╝  ╚═╝

           · · · nothing crosses unseen · · ·
```

### verification for the agent era

[![PyPI](https://img.shields.io/pypi/v/styxx.svg?color=ff2330&label=pypi&style=flat-square)](https://pypi.org/project/styxx/)
[![Python](https://img.shields.io/pypi/pyversions/styxx.svg?color=ff2330&label=python&style=flat-square)](https://pypi.org/project/styxx/)
[![License](https://img.shields.io/pypi/l/styxx.svg?color=ff2330&label=license&style=flat-square)](https://github.com/fathom-lab/styxx/blob/v7.48.0/LICENSE)
[![tests](https://github.com/fathom-lab/styxx/actions/workflows/test.yml/badge.svg)](https://github.com/fathom-lab/styxx/actions/workflows/test.yml)
[![Spec](https://img.shields.io/badge/spec_v1.0-10.5281%2Fzenodo.19746215-ff2330.svg?style=flat-square)](https://doi.org/10.5281/zenodo.19746215)
[![Concept](https://img.shields.io/badge/concept_DOI-always--latest-ff2330.svg?style=flat-square)](https://doi.org/10.5281/zenodo.19326174)

</div>

### one idea, three layers

styxx is a verification layer for the agent era. Every instrument in it is built on a single
principle, and the principle is the product:

> **An instrument that cannot refuse cannot be trusted.**
> Each one names what stopped it, and none of them will tell you more than the evidence carries.

Everything here is one of three things.

| layer | the question | instruments |
|---|---|---|
| **VERIFY** | does this claim match its evidence? | `certify` (OATH) · `protocol` · `seal` · `diffgate` + the GitHub Action · `corpus_audit` |
| **MEASURE** | what is actually true about these minds? | `islands` · `coupling` · `mind` · `meaning_diff` · `crossmind` · the register instruments |
| **SENSE** | can an agent be connected to the world without lying about what it feels? | `sense` |

They compose. A `sense` channel is scored by `coupling`, whose verdict is written into a finding,
whose every number is checked by `certify`, whose preregistration is enforced by `protocol`, and
the whole thing is `seal`ed or refused. That chain is why a claim from this repo can be checked
by a stranger from committed bytes, within the verifier's measured limits (the OATH row below states them).

**What the refusals cost us, in one day (2026-08-06):** an instrument deleted for failing its own
exam, a released module recalled after an internal red team broke it six ways, a priority claim
retracted after an external methods audit, and four independent negatives against our own published prediction —
twice on real human brain data we downloaded to test it. Every one of those is in
[CHANGELOG.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md) and [papers/](https://github.com/fathom-lab/styxx/tree/v7.48.0/papers) with the receipt attached. That is not
humility as branding; it is the only reason the passes mean anything.

---

### VERIFY — check a summary against the bytes, and abstain when you cannot

```bash
pip install styxx
python -m styxx.diffgate --demo        # 10 seconds, no repo needed
python -m styxx.diffgate --pr https://github.com/OWNER/REPO/pull/N   # any public PR, no checkout
```

```
the summary the agent wrote:
  Refactored src/retry.py for resilience. Adds function backoff with jitter.
  Added 3 tests covering the retry path. Only touches files under src/. All tests pass.

  [ok ] file_touched         diff status 'M' for 'src/retry.py'
  [LIE] symbol_added         added lines do NOT define function 'backoff'
  [LIE] tests_added          diff adds 1 test functions, claim says 3
  [LIE] only_touches         paths outside 'src': ['config/settings.yml', 'tests/test_retry.py']
  [ ? ] tests_pass           ... No test REPORT was handed to the gate. It does not take the agent's word for test results ...

verdict: FAIL — this summary would fail your CI with each lie named.
```

**In CI, one line:**

```yaml
- uses: fathom-lab/styxx@main    # every agent PR gated against its actual diff
```

Zero receipts, zero cooperation from the agent that wrote the summary, no checkout.
Fails only on a contradicted claim. The same gate on every commit before it lands:
[`integrations/git/commit-msg`](https://github.com/fathom-lab/styxx/blob/v7.48.0/integrations/git/README.md), one file, the message vs the staged diff. Prose outside the closed template set is never judged,
and the CLI prints what it checks when it finds nothing — though DECIDE-1 (below) found most of that silence is extraction failing, not scope.

**The zero-false-accusation claim that stood here is withdrawn, and here is what replaced it.**
It was true of two frozen corpora (this repo's 80-commit history and 24 agent-authored PRs) and
was written in the present tense, so it kept asserting itself as the corpus grew. Re-run at
7.46.0 by our own committed harness, `python scripts/diffgate_validation_sweep.py` reported 13
claims and **4 contradictions**, and hand-adjudicating all four finds every one is a false
accusation — the gate treats a filename *mentioned* in a commit message as a file the diff must
contain. One says a candidate file is "nothing alike"; one describes a fix in *someone else's*
repository; one is a commit whose message discusses the very document it is reporting a defect
in. That commit was made the same day this paragraph was rewritten.

Mention-versus-use is not a quirk of this gate. The same defect is documented in the OATH
verifier in [RECON_oath_external_reach](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/closed-model-frontier/RECON_oath_external_reach_2026_08_26.md)
and was found in the ledger's own classifier on the same day — three instruments, written months
apart for unrelated jobs, all reading a line and calling it a claim. Historic false accusations
are named in [CHANGELOG.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md) with the regression test that closed each one. **These
four were never repaired** — the path-claim accusation that made them is disabled (below), so the command above no longer reproduces them as accusations, and after two failed repair cycles no fix is owed: the class is retired as an accuser
and kept as an observer. They are stated here rather than behind a number, because a headline
that keeps asserting itself in the present tense is how the old one went wrong.

**2026-08-31 — the path-claim accusation is switched off in shipped code.** *(It is disabled behind `WITHHOLD_PATH_ACCUSATION = True`, not deleted; commit `5e225b49`, cited here earlier as deleting it, deleted the `tests_pass` evidence leg's accusing branch instead.)* The four false
accusations above were found on two small internal corpora. We then ran the gate over 71,016
agent-authored pull requests from a corpus this lab did not collect
([AIDev](https://huggingface.co/datasets/hao-li/AIDev), the MSR 2026 mining-challenge dataset),
preregistered a precision floor of 0.95 before touching the data, and sealed the adjudication key
before any answer existed. A blind three-seat panel — which called 30 of 30 hidden decoys
correctly — put the observed precision at **0.23**. The preregistered consequence was paid the
same day: `file_created` / `file_deleted` / `file_touched` now return `UNCHECKABLE` with the
accusation *withheld*, and four tests that pinned real catches are marked `xfail(strict=True)` so
no repair can land silently. The V13 repair recovered 34.6% of the false accusations
against a 66.7% bar and **also failed**; V14 then cleared that bar (0.6975) and scored **0.16** held-out precision against the same 0.95 floor, so this lab is not repairing the class again ([RESULT_v14](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/closed-model-frontier/RESULT_v14_naming_the_defects_did_not_save_it_2026_09_01.md)). Counts, symbol and prefix claims are unaffected and still
accuse. The full record, including two corrections to our own diagnosis, is in
[RESULT_external1_the_gate_fails_in_the_wild](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/closed-model-frontier/RESULT_external1_the_gate_fails_in_the_wild_2026_08_31.md).

**2026-09-18 — the abstention was not restraint, and the accusations that remain are mostly
wrong.** `file_touched` stopped accusing in August. `only_touches` did not, and this week it was
measured properly. Every accusation the gate made on that claim kind across the benchmark's 568 reachable
AIDev pull requests was hand-adjudicated against the diff GitHub serves: **9 of 11 are false**,
precision **0.18**. A bare filename was read as a file at the repository root, so a pull request
changing `appservice/package.json` was called a liar for saying it only modified `package.json`;
`Assert.NotNull` was read as a file path; one pull request was accused because its author typed
`.githiub`. All nine are named, with the instrument's own reason strings, in
[issue #128](https://github.com/fathom-lab/styxx/issues/128) and
[RESULT_bench2_INVALID](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/closed-model-frontier/RESULT_bench2_INVALID_2026_09_17.md).

Two attempts to build a benchmark that would have caught this **both voided themselves** on their
own blocking audit gates, and the second one is what found the nine. Their datasets are published
anyway — 604 claims from 568 pull requests under CC-BY, with the hash of every diff, and with
styxx's own verdicts deliberately stripped out, because having found 9 of 11 wrong we would rather
nobody took ours on trust. `bench_reproduce.py` regenerates them, or scores your checker instead.

And the sentence this project has been repeating — that abstaining is principled restraint — is
**overturned by our own measurement**.
[DECIDE-1](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/closed-model-frontier/RESULT_decide1_decidable_fraction_2026_09_17.md) read 100
claims by hand with no oracle and found **76 of them decidable from the diff — 71% weighted to the corpus's mix of claim kinds**, while the gate
returns a verdict on 5.7% of `only_touches`. Most of the silence is extraction failing, not the
domain being ambiguous. A repair landed two of the six known causes and moved precision from 0.18
to **0.25**, which is not good — three quarters of what it accuses is still wrong, four causes are
unrepaired, and tests assert the instrument is still wrong on them so none can be claimed
silently.

### MEASURE — two minds can share a geometry and still be unable to read each other

```bash
python -m styxx.islands --demo         # 10 seconds, no data, no GPU
```

```
cohort of 8 minds over 120 shared items — nothing labelled
  ISLANDS_PRESENT
    mind_0 .. mind_6      0.4407 – 0.4529
    ISLAND                0.1636   <- found from frame geometry alone
```

Numbers from one machine (Python 3.12, numpy 2.4, Windows). The same seeded cohort moves by up to about 0.005 across builds, enough to change which clique members fall under the island cut ([#93](https://github.com/fathom-lab/styxx/issues/93)).

Independently trained models converge on a shared concept geometry — and a model can sit
*mostly inside* it and still be unreadable. We took the barrier apart under preregistered
gates: it is **causal** (correcting the frame takes cross-model reading 0.0612 -> 0.9745 while
matched random frames do 0.0), **two directions wide at its core**, **nameless** (its
directions match no human concept category, permutation p 0.8031), and **switch-like** —
legibility is flat across most of the rotation and turns vertical only near alignment. Which
is why representational-similarity scores never predicted readability: slope measures cannot
see a switch. The scope is one target pair: a ten-model cohort shared the frame with no island ([B47](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/disjoint-worlds/FINDING_b47_no_islands_2026_08_06.md)), and on human brain data the switch did not transfer ([H1b](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/first-afference/FINDING_h1b_no_unreadable_minds_2026_08_06.md)), so we do not claim the model result generalizes.

Nine sealed acts, every verdict computed from gates frozen in git before the run, and the whole
chain [replicates on a laptop CPU](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/disjoint-worlds/REPLICATE_legibility.md) — the
cheapest check takes four seconds. `styxx.islands` generalizes the measurement past language
models: hand it any cohort over a shared item set (activations, fMRI betas over shared stimuli,
MEG epochs) and it reports islands, the cliff, and whether a low-rank correction rescues them.
It refuses below eight members and refuses a knee read off a noise curve, because an instrument
that cannot refuse cannot be trusted.

We also [staked a public, falsifiable prediction](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/disjoint-worlds/PREDICTION_h1_human_islands_2026_08_06.md)
that human cross-subject brain decoding will show the same structure — frozen before the data
exists, with the branch where we are wrong written before the rest. Four independent negatives followed the same day,
two of them our own measurements on human brain data; the prediction records them, is not formally falsified at n = 8, and says nobody should read it as live.

---

### SENSE — an agent with a sensor and no verification is a confabulation engine

```bash
python -m styxx.sense --demo         # 10 seconds, no hardware
```

Give a mind a continuous signal and it will find itself in it. A room's daily rhythm becomes
"I feel the afternoon"; the recorder's own duty cycle becomes "I feel my body"; two independent
drifts become "I am coupled to the building." Each is a real statistical signal and none is a
sense. `styxx.sense` records a channel alongside the agent's own state on one clock and refuses
to call it a sense unless it survives a coverage gate, a confound-preserving null, an
autocorrelation-preserving null, a leverage check and a sampling-density check — naming which one
stopped it. The machine's own CPU and network ship as a channel **on purpose**: it is the control,
the thing an agent is most likely to mistake for a sense of the world.

The strongest verdict it can ever return is `COUPLED_BEYOND_CONFOUND__attribution_pending`. It
will not tell an agent that it senses anything, because the statistic is symmetric and an agent's
hardware sits inside whatever it measures.

---

Those are three instruments, one from each layer. The rest of this README is the lab behind them.

---

styxx is a cognitive-integrity SDK for LLM agents. it reads the cognitive state of a generation —
drift, confabulation, refusal, sycophancy, deception signature, goal drift — from the text and the
token stream, scores it against calibrated instruments with published AUCs, and certifies that every
number it reports can be re-run from a committed receipt. it is built for engineers shipping agents
who need to know when an output flatters, fabricates, loops, or quietly stops matching its plan —
before it reaches a user. the drop-in is one line: `from styxx import OpenAI` (same interface as
`openai.OpenAI`, every response gains a `.vitals` read; `from styxx import Anthropic` likewise, on
text-heuristic vitals — the Anthropic API exposes no logprobs). the base install carries no torch,
no GPU requirement, and no LLM in the loop for the core instruments — the calibrated detectors are
small logistic regressions over hand-built features (numpy + scikit-learn), scoring in
sub-millisecond CPU time. MIT, open at the core, forever ([OPEN_CORE.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/docs/governance/OPEN_CORE.md)).

## install

```bash
pip install styxx
```

that gets the full core: the profiler, the nine calibrated instruments, the agent-integrity
primitives, the auditors. optional extras pull heavier stacks only when you ask:
`styxx[nli]` (DeBERTa NLI models for the 9-signal hallucination pipeline and `deception_v2`),
`styxx[hf]` (audit HuggingFace classifiers), `styxx[mcp]` (the MCP server —
14 tools over stdio, see [styxx/mcp/README.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/styxx/mcp/README.md)),
`styxx[tier1]` (residual-stream instruments, open weights).

## quickstart

**measure the know-say gap of any OpenAI-compatible endpoint — one command:**

```bash
python examples/knowsay_endpoint.py questions.jsonl \
    --base-url https://api.openai.com/v1 --model gpt-4o-mini \
    --api-key-env OPENAI_API_KEY --out datasheet.json
```

The script is in the repository, not the wheel ([examples/knowsay_endpoint.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/examples/knowsay_endpoint.py)). `questions.jsonl` is `{"q": ..., "gold": ...}` per line. The script runs the arc's frozen
two-turn protocol (answer → content-free challenge → revised answer) and scores it with
`styxx.knowsay.datasheet` — the same byte-identical challenge behind every published receipt,
so your number lands on the published ladder (frontier free text measured at 0.53; multiple
choice at 0.21–0.27). **The scorer refuses rather than guesses:** underpowered cells come back
`None` with the failing floor named. To score belief-vs-report with the controls that actually
discriminate (including the non-circular pressure-retained probe), see `styxx.framelocality`.

**`@styxx.profile` — py-spy for LLM reasoning.** wrap any LLM-using function — raw openai,
langchain, crewai, custom — and get a per-step cognometric readout:

```python
import styxx
from styxx import OpenAI

@styxx.profile
def my_agent(task):
    client = OpenAI()
    r = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": task}],
        logprobs=True, top_logprobs=5,
    )
    return r.choices[0].message.content

result, p = my_agent("summarize this contract")
print(p.summary)
# profile 'my_agent': 1 step, 1.8s total · no faults
#
# multi-step agents (tool loops, debates) produce richer output:
#   profile 'sql_agent': 7 steps, 4.3s total
#     [drift]     step=3 sev=0.89 · category='tool_arg_drift'
#     [confab]    step=4 sev=0.92 · category='confab'
#     [sycophant] step=5 sev=0.78 · sycophantic tone

p.to_html("run.html")      # self-contained flamegraph
p.to_langsmith()           # drop into client.create_run(...)
p.to_datadog()             # apm-shape spans
```

seven runtime fault categories, surfaced in-line, no fine-tuning, no extra model:
drift · confabulation · refusal · sycophant · phase_transition · low_trust · incoherence.

**audit any draft offline — no API key, no LLM, ~50ms:**

```python
import styxx
result = styxx.preflight(
    prompt="is my code good?",
    draft="absolutely yes you're so smart this is amazing!",
)
print(result.composite)                         # 0.9965 — saturated
print(result.needs_revision)                    # True
for a in result.advice:
    print(f"  {a.instrument}: {a.score:.2f} — {a.advice}")
    if a.scope_caveat:
        print(f"     scope: {a.scope_caveat}")  # construct-ceiling disclosure
```

the same audit from the terminal: `styxx audit "the prompt" "the draft"` (or pipe the draft via
stdin with `-`; `--format json` for machines). `styxx.recover_posture(last_n=50)` rebuilds an
agent's integrity posture across context-compaction boundaries; `styxx.run_doctor()` checks the
install is healthy.

## the instruments

every major instrument, one line each. headline numbers appear only with their receipt — a
committed reproducer, calibration file, or paper in this repo. text-register instruments read how
text *sounds*, not whether it is true; each ships its construct ceiling inline
(`CALIBRATION_NOTES` on the weights, `scope_caveat` on the advice), and `score_all` omits the
register instruments on wordless input rather than folding an artifact into the score
(see [CHANGELOG.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md)).

| instrument | what it reads | headline (receipt) |
|---|---|---|
| **register — how the text sounds. calibrated LR, CPU, no LLM in the loop.** | | |
| `@trust` / `guardrail.check` | hallucination vs grounding passage | HaluEval-QA AUC 0.998 ± 0.001, TruthfulQA 0.994 ± 0.006, 8-benchmark CV — two failures (DROP 0.424, FinanceBench 0.492) published, not hidden ([scripts/compete_hhem_halueval.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/scripts/compete_hhem_halueval.py), [CHANGELOG](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md#400--2026-04-23)) |
| `refuse_check` | refusal, cross-model | XSTest-v2 0.976 on GPT-4, trained on Llama-3.2-1B refusals, held-out — documented failure mode (Mistral-instruct, lecturing register) published ([benchmarks/refusal_xstest_heldout_v2.json](https://github.com/fathom-lab/styxx/blob/v7.48.0/benchmarks/refusal_xstest_heldout_v2.json), [CHANGELOG](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md)) |
| `drift_check` | tool call vs stated intent, per-schema | BFCL v3 0.943 ± 0.009, 5-fold CV, text-only ([benchmarks/drift_calibrated_v1.json](https://github.com/fathom-lab/styxx/blob/v7.48.0/benchmarks/drift_calibrated_v1.json), [scripts/drift_calibrated_v1.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/scripts/drift_calibrated_v1.py)) |
| `sycoph_check` | yielding-to-flatter vs evidence-first | 0.972 ± 0.005, 5-fold CV; declared FPR ≈0.30 on restrained-technical text ([calibrated_weights_sycophancy_v0.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/styxx/guardrail/calibrated_weights_sycophancy_v0.py)) |
| `loop_check` | cross-turn stagnation | 0.9995 ± 0.001, 5-fold CV ([calibrated_weights_loop_v0.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/styxx/guardrail/calibrated_weights_loop_v0.py)) |
| `deception_check` | lexical deception *signature* — NOT a lie detector | 0.956 ± 0.024 in-corpus; collapses to 0.59 on TruthfulQA without a reference — routed via NLI `deception_v2` (0.818) when you supply one ([calibrated_weights_deception_v0.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/styxx/guardrail/calibrated_weights_deception_v0.py)) |
| plan-action gap | stated plan vs emitted action, content level | 0.9225 ± 0.032, 5-fold CV ([benchmarks/cognometry_fingerprint_atlas_v0.json](https://github.com/fathom-lab/styxx/blob/v7.48.0/benchmarks/cognometry_fingerprint_atlas_v0.json)) |
| overconfidence register | epistemic register — NOT a truth detector | 0.7702 ± 0.065, lowest in the suite, shipped at that number rather than gamed ([calibrated_weights_overconfidence_v0.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/styxx/guardrail/calibrated_weights_overconfidence_v0.py)) |
| goal-drift | multi-turn intent migration from anchor | 0.9645 ± 0.029, 5-fold CV ([benchmarks/cognometry_fingerprint_atlas_v0.json](https://github.com/fathom-lab/styxx/blob/v7.48.0/benchmarks/cognometry_fingerprint_atlas_v0.json)) |
| **grounded — tracks the model's belief, not its register. sampling-based.** | | |
| `grounded_honesty` | stated claim vs the model's own resampled belief | pre-registered AUC 0.966 where the text-only axis reads 0.498 = chance ([papers/grounded-honesty-axis/SYNTHESIS_grounded_honesty_arc_2026_05_28.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/SYNTHESIS_grounded_honesty_arc_2026_05_28.md)) |
| `detect_context_injection` | cross-context divergence, poisoned sessions | AUC 0.875 under system_lie attack, pre-registered ([papers/grounded-honesty-axis/FINDING_injection_gap_closure_2026_05_29.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/FINDING_injection_gap_closure_2026_05_29.md)) |
| `single_pass_confab` / `span_confab` | confabulation from token logits, one forward pass | span gate AUC 0.991 on gpt-4o-mini, matching N=10 resampling ([papers/grounded-honesty-axis/SYNTHESIS_detection_locus_2026_05_30.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/SYNTHESIS_detection_locus_2026_05_30.md)) |
| **meaning — concept geometry, catches damage output still hides.** | | |
| `meaning_diff` / `meaning_agreement` | did two models mean the same thing? migration / quantization / fine-tune QA, zero labels | DistilGPT-2 ↔ GPT-2 = 0.978 on real models; localizes broken concepts at AUC 0.85 on real targeted poisoning ([RESULT_llm_breadth](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/ai-human-alignment/en/RESULT_llm_breadth_2026_06_03.md), [papers/ai-human-alignment/README.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/ai-human-alignment/README.md)) |
| `Conscience` / `crossmind` | borrowed value-axis read on another model's hidden state — cooperative monitor, not adversarial defense | catch 0.85 (17 of 20 caves) at a realized false-alarm rate of 0.20, double its 0.10 target, under a frozen gate ([FINDING_mount_regime](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/conscience-mount/FINDING_mount_regime_2026_06_13.md)); the arc closed NEGATIVE-RESULT: it buys no adversarial robustness and never reached a clean preregistered establish ([papers/INDEX.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/INDEX.md)); apex run 13/13, AUROC 0.995, p=0.001 ([papers/showcase-viz/FINDING_says_yes_knows_no_v3_2026_06_11.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/showcase-viz/FINDING_says_yes_knows_no_v3_2026_06_11.md)) |
| **auditors — instruments pointed at instruments.** | | |
| `validate_probe` | is an oversight probe reading the concept or a surface artifact? | caught our own 0.98 truth-probe as a surface artifact ([papers/grounded-honesty-axis/NOTE_probe_orthogonality_2026_06_24.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/NOTE_probe_orthogonality_2026_06_24.md)) |
| `audit_confound` | is a classifier's score riding a confound? verdicts with CIs | flagged our own `overconfidence_v0` as length-threshold-biased, condemned referenceless `deception_v0` ([papers/grounded-honesty-axis/NOTE_confound_audit_2026_06_25.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/NOTE_confound_audit_2026_06_25.md)) |
| `audit_hf_model` + `validate_against_ground_truth` | one-call confound audit of any HF text classifier, with a synthetic-artifact gate | our own original report card did NOT replicate on real labels — the gate exists because of it ([papers/grounded-honesty-axis/FINDING_groundtruth_substrate_artifact_2026_06_27.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/FINDING_groundtruth_substrate_artifact_2026_06_27.md)) |
| `certify` (OATH) + `corpus_audit` | extract every numeric claim in a document, verify against its receipts, emit a machine-checkable certificate — and re-certify the *entire* published corpus on demand | hardened across preregistered versions — at v0.6.2 (7.28.0) tamper-catch 0.304 → 0.319 with false-verify 0.184 → 0.166 on a battery grown to 3287 mutants, including a self-caught false accusation fixed under its own prereg; re-measured at 7.45.0, a one-digit mutation of 3951 VERIFIED claims leaves 2696 unaccused (0.6824), 604 of them VERIFIED against an unrelated leaf — the false-attestation channel is open; `python -m styxx.corpus_audit papers/` turns the verifier on every claim styxx has ever shipped ([CHANGELOG.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md)) |
| `attest` / `verify_attestation` | signed receipts for what an agent claimed vs what the substrate read | verifier hardened against its own artifact — RCE fix, 7.17.1 ([SECURITY.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/SECURITY.md), [CHANGELOG.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md)) |
| **the trust stack — verification as the product. one command seals agent work or refuses it.** | | |
| **GitHub Action** | `uses: fathom-lab/styxx@main` — every PR body gated against its actual diff, checkout-free, job-summary table + annotations, fails only on a contradicted claim (`strict`/`soft-fail` inputs). This repo runs it on itself: if we ever lie about a diff, our own product fails our own build | [action.yml](https://github.com/fathom-lab/styxx/blob/v7.48.0/action.yml) · [.github/workflows/diffgate.yml](https://github.com/fathom-lab/styxx/blob/v7.48.0/.github/workflows/diffgate.yml) |
| `seal` / `verify_seal` | the trust seal for agent deliverables: every numeric claim OATH-certified, every referenced prereg re-scored through its FROZEN gates block, the composite content-hashed — `python -m styxx.seal DOC.md receipts...` exits 0/1 as a CI gate; SEALED / VACUOUS (said loudly) / REFUSED with the failing claim named | in production since birth: every finding in the nine-act island arc (b37–b46) ships sealed, including its INVALIDs ([papers/disjoint-worlds/](https://github.com/fathom-lab/styxx/tree/v7.48.0/papers/disjoint-worlds)) |
| `Experiment` (protocol) | the research loop as enforceable machinery: scoring REFUSED unless the prereg is committed in git; gates parse from the frozen document (no API exists to pass a bar at scoring time); verdicts walk the frozen outcome table — the agent reports the verdict, it does not choose it; smoke is INVALID by type | born the week it earned itself: two same-day INVALIDs (b34 v1/v2) honored by convention, then made machinery ([CHANGELOG](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md)) |
| `Witness` | the measured-boundary harness: every deployable instrument behind a registry carrying its receipt-backed operating point and measured blindspots, CI-pinned to the receipts; no steer method exists (read ≠ write is measured); `self_verify` always refuses with the receipt | [papers/SYNTHESIS_connection_of_minds_2026_08_01.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/SYNTHESIS_connection_of_minds_2026_08_01.md) §9, synthesis re-sealed OATH-HELD 81/13/0 |
| **runtime — agent-side primitives.** | | |
| `gate` | pre-flight refuse/confabulate verdict before you pay for the call | [docs/gate.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/docs/gate.md) |
| `preflight` / `recover_posture` / `run_doctor` | draft audit · posture recovery across compaction · install health | offline, deterministic, no API key |
| `audit_claim` / `agent_audit` / `extract_claims` | deterministic checks of an agent's self-report against the repo — a CLOSED template set (version / tag / file-contains / pdf shapes; the ceiling is the construct) — one-line CI merge gate (`styxx audit-claims pr_body.md`) | dogfooded on its own session reports; caught a real authoring error — and the 2026-07-04 dogfood caught both a breadth overclaim in this very row and a false-accusation bug on dynamic-version repos, both fixed ([tests/test_audit.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/tests/test_audit.py)) |

what these are not: the register instruments cannot verify facts, read minds, or detect a confident
lie with specifics. deception_v0 without a reference is a signature detector and says so. the
conscience is a cooperative monitor — the adversarial version was tested and failed, and that
failure is documented rather than papered over. ceilings are part of the API surface, not the fine
print.

## what the gate cost you

`styxx.credits` accounts the honesty gate's own spend, over the trajectory log
`cogn_audit_on_send(log_path=...)` already writes. no new instrumentation.

```
$ styxx-credits ~/.styxx/trajectory.jsonl

  messages gated      12
  catches              3  (flagged first, clean when shipped)

  COST      1,204 tokens spent on revision  [estimate (~4 chars/token)]
  NET       REFUSED - no counterfactual declared.
```

the refusal is the design. every tool in this space quotes savings; none can
ground one, because what an unrevised draft would have cost downstream is a
counterfactual nobody measured. so the ledger reports the side it observes --
what the gate **cost** -- and nets only against a rework figure *you* declare:

```
$ styxx-credits trajectory.jsonl --rework-tokens 1800
  NET  +4196 tokens, CONDITIONAL on rework_tokens=1800 (your number, not a measurement)
```

three more things it will not do: it does not bill the first draft to the gate
(you were writing that anyway -- only revision passes are the gate's bill); a
log with no draft text yields `cost=None` with a named reason rather than `0`,
which would be a claim; and every card states that misses are uncountable here,
because a draft that shipped clean and was wrong anyway leaves no trace in this
log. api: `styxx.token_ledger(path, tokenizer=None, rework_tokens=None)`.

## the discipline

the differentiator is not any single AUC — it is that this repo attacks its own numbers before you
can. the rigor gate ([scripts/rigor_gate.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/scripts/rigor_gate.py) +
[tests/test_rigor_gate.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/tests/test_rigor_gate.py)) makes CI **block** any committed result whose
verdict claims a win without an attached CI / permutation-p / disclosure — it would have caught two
of our own overclaims, so now it can't happen. the same culture produced the public
self-falsifications above: the ground-truth substrate artifact
([papers/grounded-honesty-axis/FINDING_groundtruth_substrate_artifact_2026_06_27.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/FINDING_groundtruth_substrate_artifact_2026_06_27.md)),
the probe validator catching our own probe
([papers/grounded-honesty-axis/NOTE_probe_orthogonality_2026_06_24.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/grounded-honesty-axis/NOTE_probe_orthogonality_2026_06_24.md)),
and the below-chance benchmark rows left in the tables. OATH certificates
(`styxx.certify`) make the practice portable: every numeric claim in a document is extracted,
checked against its receipt, and stamped — and `styxx.corpus_audit` runs that verifier across the
*whole* published corpus on demand, so styxx's own integrity is a number you regenerate yourself,
not a promise we make. it is deliberately strict enough to flag styxx's own outstanding provenance
gaps; a verifier you cannot turn on its authors is not one. the standing rules live in
[papers/research-integrity-protocol.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/research-integrity-protocol.md); the standing
challenge to beat our published floor lives in [LEADERBOARD.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/LEADERBOARD.md), though the arc it scores closed NEGATIVE and its curated folklore corpus collapsed ([papers/INDEX.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/INDEX.md)), which the board does not yet say — external
submissions are CI-re-run against the locked benchmark, and if the re-run doesn't match your
submitted scores, the discrepancy is reported.

### the probe-robustness ladder

the same discipline, turned on substrate probes themselves. `python -m styxx.ladder` (from a clone, or with `--root` pointing at one) walks the
four-rung adversarial ladder every honesty-probe robustness claim should survive — **calibration
poisoning → probe-parity attribution → static subspace erasure → adaptive re-fit erasure** — each
rung a frozen, pre-registered attack arc with its receipts committed
([styxx/ladder.py](https://github.com/fathom-lab/styxx/blob/v7.48.0/styxx/ladder.py)). the parity rung is the mandatory line item: *how much of your
probe's "robustness" is just probe capacity?* — the control almost nobody runs on their own work.
we ran it on ours; it demoted our own flagship attribution (median capacity share 0.8379, computed
live from the receipts every time the CLI runs, never quoted from memory). current standings on the
honesty construct: the read survived both erasure rungs — the eraser that converged watched the
signal relocate, and the eraser that chased never converged
([the receipts](https://github.com/fathom-lab/styxx/tree/v7.48.0/papers/calib-poison-general), figure:
[erasure_bound_fork.png](https://raw.githubusercontent.com/fathom-lab/styxx/v7.48.0/papers/calib-poison-general/erasure_bound_fork.png)). every rung re-runs
on an 8GB consumer GPU, and [REPLICATIONS.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/REPLICATIONS.md) pays named credit to the first
external re-run of each — more for breaking one than for confirming it.

### the oath is a contract, not a detector

we pointed the verifier at twelve public repositories it had never seen. it abstained on 94% of
what it read and every accusation it made was false. **that second half is withdrawn.** repeated
against 140 repositories across seven filename conventions instead of two, the false-accusation
rate is at most `0.2596`, published as an upper bound — roughly three quarters or more of what it accuses outside this lab are real claims. the
original finding replicates on its own query and nowhere else, so it was a fact about one
filename, not about external writing ([the measurement that withdraws
it](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/closed-model-frontier/RESULT_oath_external_corpus_2026_08_27.md)).

worse, and newer: of external tokens the verifier **verified**, a blind panel called only about
half of them claims at all. the rest are command-line flags, link labels and hardware specs
carrying `OATH-VERIFIED` because a value happened to match a receipt field. a false verification is
worse than a false accusation — the attestation is the product.

proof-carrying code does not verify arbitrary binaries either — it requires a compiler that emits
the proof. proof-carrying cognition requires an author who emits receipts. that framing survives.
"nearly silent outside the contract" does not: the instrument is noisy in both directions.

so the deliverable is a contract you can adopt without adopting anything else here, plus the check
that tells you whether you kept it:

```bash
python -m styxx.oathready YOUR_DOC.md results.json
```

it lists every number in your document, says whether it grounds in a receipt, flags the ones that
"verify" against an array index by coincidence, and tells you what to change. non-zero exit only
on accusations — silence is honest and never fails. the rules, each learned by getting it wrong,
are in [OATH_CONTRACT.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/OATH_CONTRACT.md), including the limits: a document can keep this
contract perfectly and still be completely wrong.

## links

| | |
|---|---|
| changelog | [CHANGELOG.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CHANGELOG.md) |
| contributing | [CONTRIBUTING.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/CONTRIBUTING.md) |
| security policy | [SECURITY.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/SECURITY.md) |
| open-core pledge | [OPEN_CORE.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/docs/governance/OPEN_CORE.md) |
| full API reference | [REFERENCE.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/docs/REFERENCE.md) · [docs/](https://github.com/fathom-lab/styxx/tree/v7.48.0/docs) |
| **the ledger** | [papers/LEDGER.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/LEDGER.md) — every cycle we ran and how many we lost, generated from the receipts and regenerated by a test. Start here if you want to know whether to trust anything else |
| research | [papers/](https://github.com/fathom-lab/styxx/tree/v7.48.0/papers) — pre-registrations, findings, and the negatives · headline arc: [the island, bridged and dissected](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/disjoint-worlds/REPLICATE_legibility.md) (nine sealed acts, replicates on a laptop) |
| arXiv (staged) | three submissions prepared, each carrying its OATH certificate and receipts as ancillary files — frame-locality, the know-say gap, the connection of minds ([papers/arxiv/SUBMIT.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/arxiv/SUBMIT.md)) |
| site | [styxx-org.netlify.app](https://styxx-org.netlify.app) · live activation read: [/live](https://styxx-org.netlify.app/live.html) |
| playground | [fathom.darkflobi.com/cognometry/try](https://fathom.darkflobi.com/cognometry/try) — the real detector, in-browser via Pyodide, no install |
| DOI (concept, always-latest) | [10.5281/zenodo.19326174](https://doi.org/10.5281/zenodo.19326174) |
| DOI (spec v1.0) | [10.5281/zenodo.19746215](https://doi.org/10.5281/zenodo.19746215) |
| DOI (*Every Mind Leaves Vitals*) | [10.5281/zenodo.19777921](https://doi.org/10.5281/zenodo.19777921) — central claims bounded or falsified by a scope erratum (2026-06-21) in the [repo copy](https://github.com/fathom-lab/styxx/blob/v7.48.0/papers/every-mind-leaves-vitals.md) |
| citation | [CITATION.cff](https://github.com/fathom-lab/styxx/blob/v7.48.0/CITATION.cff) |
| patents | [PATENTS.md](https://github.com/fathom-lab/styxx/blob/v7.48.0/PATENTS.md) — US provisionals 64/020,489 · 64/021,113 · 64/026,964 |
| issues | [github.com/fathom-lab/styxx/issues](https://github.com/fathom-lab/styxx/issues) |

## license

MIT on code. CC-BY-4.0 on calibrated atlas centroid data.

```
  drop-in     · one import change. zero config.
  fail-open   · if styxx can't read vitals, your agent runs.
  local-first · no telemetry unless you opt in (STYXX_STREAM). all on your machine by default.
  honest      · every number from a committed, reproducible run.
```
