Metadata-Version: 2.5
Name: trust-but-anchor
Version: 0.1.0
Summary: The model proposes a short anchor; code locates a real source span. Fail closed.
Project-URL: Homepage, https://github.com/gmhoward9289-ops/trust-but-anchor
Author-email: "George M. Howard" <dev@swamplink.com>
License:                                  Apache License
                                   Version 2.0, January 2004
                                http://www.apache.org/licenses/
        
           TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
        
           1. Definitions.
        
              "License" shall mean the terms and conditions for use, reproduction,
              and distribution as defined by Sections 1 through 9 of this document.
        
              "Licensor" shall mean the copyright owner or entity authorized by
              the copyright owner that is granting the License.
        
              "Legal Entity" shall mean the union of the acting entity and all
              other entities that control, are controlled by, or are under common
              control with that entity. For the purposes of this definition,
              "control" means (i) the power, direct or indirect, to cause the
              direction or management of such entity, whether by contract or
              otherwise, or (ii) ownership of fifty percent (50%) or more of the
              outstanding shares, or (iii) beneficial ownership of such entity.
        
              "You" (or "Your") shall mean an individual or Legal Entity
              exercising permissions granted by this License.
        
              "Source" form shall mean the preferred form for making modifications,
              including but not limited to software source code, documentation
              source, and configuration files.
        
              "Object" form shall mean any form resulting from mechanical
              transformation or translation of a Source form, including but
              not limited to compiled object code, generated documentation,
              and conversions to other media types.
        
              "Work" shall mean the work of authorship, whether in Source or
              Object form, made available under the License, as indicated by a
              copyright notice that is included in or attached to the work
              (an example is provided in the Appendix below).
        
              "Derivative Works" shall mean any work, whether in Source or Object
              form, that is based on (or derived from) the Work and for which the
              editorial revisions, annotations, elaborations, or other modifications
              represent, as a whole, an original work of authorship. For the
              purposes of this License, Derivative Works shall not include works
              that remain separable from, or merely link (or bind by name) to the
              interfaces of, the Work and Derivative Works thereof.
        
              "Contribution" shall mean any work of authorship, including
              the original version of the Work and any modifications or additions
              to that Work or Derivative Works thereof, that is intentionally
              submitted to Licensor for inclusion in the Work by the copyright owner
              or by an individual or Legal Entity authorized to submit on behalf of
              the copyright owner. For the purposes of this definition, "submitted"
              means any form of electronic, verbal, or written communication sent
              to the Licensor or its representatives, including but not limited to
              communication on electronic mailing lists, source code control systems,
              and issue tracking systems that are managed by, or on behalf of, the
              Licensor for the purpose of discussing and improving the Work, but
              excluding communication that is conspicuously marked or otherwise
              designated in writing by the copyright owner as "Not a Contribution."
        
              "Contributor" shall mean Licensor and any individual or Legal Entity
              on behalf of whom a Contribution has been received by Licensor and
              subsequently incorporated within the Work.
        
           2. Grant of Copyright License. Subject to the terms and conditions of
              this License, each Contributor hereby grants to You a perpetual,
              worldwide, non-exclusive, no-charge, royalty-free, irrevocable
              copyright license to reproduce, prepare Derivative Works of,
              publicly display, publicly perform, sublicense, and distribute the
              Work and such Derivative Works in Source or Object form.
        
           3. Grant of Patent License. Subject to the terms and conditions of
              this License, each Contributor hereby grants to You a perpetual,
              worldwide, non-exclusive, no-charge, royalty-free, irrevocable
              (except as stated in this section) patent license to make, have made,
              use, offer to sell, sell, import, and otherwise transfer the Work,
              where such license applies only to those patent claims licensable
              by such Contributor that are necessarily infringed by their
              Contribution(s) alone or by combination of their Contribution(s)
              with the Work to which such Contribution(s) was submitted. If You
              institute patent litigation against any entity (including a
              cross-claim or counterclaim in a lawsuit) alleging that the Work
              or a Contribution incorporated within the Work constitutes direct
              or contributory patent infringement, then any patent licenses
              granted to You under this License for that Work shall terminate
              as of the date such litigation is filed.
        
           4. Redistribution. You may reproduce and distribute copies of the
              Work or Derivative Works thereof in any medium, with or without
              modifications, and in Source or Object form, provided that You
              meet the following conditions:
        
              (a) You must give any other recipients of the Work or
                  Derivative Works a copy of this License; and
        
              (b) You must cause any modified files to carry prominent notices
                  stating that You changed the files; and
        
              (c) You must retain, in the Source form of any Derivative Works
                  that You distribute, all copyright, patent, trademark, and
                  attribution notices from the Source form of the Work,
                  excluding those notices that do not pertain to any part of
                  the Derivative Works; and
        
              (d) If the Work includes a "NOTICE" text file as part of its
                  distribution, then any Derivative Works that You distribute must
                  include a readable copy of the attribution notices contained
                  within such NOTICE file, excluding those notices that do not
                  pertain to any part of the Derivative Works, in at least one
                  of the following places: within a NOTICE text file distributed
                  as part of the Derivative Works; within the Source form or
                  documentation, if provided along with the Derivative Works; or,
                  within a display generated by the Derivative Works, if and
                  wherever such third-party notices normally appear. The contents
                  of the NOTICE file are for informational purposes only and
                  do not modify the License. You may add Your own attribution
                  notices within Derivative Works that You distribute, alongside
                  or as an addendum to the NOTICE text from the Work, provided
                  that such additional attribution notices cannot be construed
                  as modifying the License.
        
              You may add Your own copyright statement to Your modifications and
              may provide additional or different license terms and conditions
              for use, reproduction, or distribution of Your modifications, or
              for any such Derivative Works as a whole, provided Your use,
              reproduction, and distribution of the Work otherwise complies with
              the conditions stated in this License.
        
           5. Submission of Contributions. Unless You explicitly state otherwise,
              any Contribution intentionally submitted for inclusion in the Work
              by You to the Licensor shall be under the terms and conditions of
              this License, without any additional terms or conditions.
              Notwithstanding the above, nothing herein shall supersede or modify
              the terms of any separate license agreement you may have executed
              with Licensor regarding such Contributions.
        
           6. Trademarks. This License does not grant permission to use the trade
              names, trademarks, service marks, or product names of the Licensor,
              except as required for reasonable and customary use in describing the
              origin of the Work and reproducing the content of the NOTICE file.
        
           7. Disclaimer of Warranty. Unless required by applicable law or
              agreed to in writing, Licensor provides the Work (and each
              Contributor provides its Contributions) on an "AS IS" BASIS,
              WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
              implied, including, without limitation, any warranties or conditions
              of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
              PARTICULAR PURPOSE. You are solely responsible for determining the
              appropriateness of using or redistributing the Work and assume any
              risks associated with Your exercise of permissions under this License.
        
           8. Limitation of Liability. In no event and under no legal theory,
              whether in tort (including negligence), contract, or otherwise,
              unless required by applicable law (such as deliberate and grossly
              negligent acts) or agreed to in writing, shall any Contributor be
              liable to You for damages, including any direct, indirect, special,
              incidental, or consequential damages of any character arising as a
              result of this License or out of the use or inability to use the
              Work (including but not limited to damages for loss of goodwill,
              work stoppage, computer failure or malfunction, or any and all
              other commercial damages or losses), even if such Contributor
              has been advised of the possibility of such damages.
        
           9. Accepting Warranty or Additional Liability. While redistributing
              the Work or Derivative Works thereof, You may choose to offer,
              and charge a fee for, acceptance of support, warranty, indemnity,
              or other liability obligations and/or rights consistent with this
              License. However, in accepting such obligations, You may act only
              on Your own behalf and on Your sole responsibility, not on behalf
              of any other Contributor, and only if You agree to indemnify,
              defend, and hold each Contributor harmless for any liability
              incurred by, or claims asserted against, such Contributor by reason
              of your accepting any such warranty or additional liability.
        
           END OF TERMS AND CONDITIONS
        
           Copyright 2026 George M. Howard
        
           Licensed under the Apache License, Version 2.0 (the "License");
           you may not use this file except in compliance with the License.
           You may obtain a copy of the License at
        
               http://www.apache.org/licenses/LICENSE-2.0
        
           Unless required by applicable law or agreed to in writing, software
           distributed under the License is distributed on an "AS IS" BASIS,
           WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
           See the License for the specific language governing permissions and
           limitations under the License.
License-File: LICENSE
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == 'dev'
Description-Content-Type: text/markdown

# Trust, But Anchor

[![Discussions](https://img.shields.io/github/discussions/gmhoward9289-ops/trust-but-anchor)](https://github.com/gmhoward9289-ops/trust-but-anchor/discussions)

**On adversarial documents, models produce a character-for-character verbatim quote only 64–73% of the time — even when explicitly told to.** Letting code locate a short model-proposed anchor phrase instead recovers 91–100% coverage, with every emitted span guaranteed to be a real substring of the source. That's the whole argument: don't trust the model's quote — trust its anchor, and verify the anchor in code.

**52 local models measured**, and only 11 clear a usable bar — ≥95% verified coverage *and* a perfect record of refusing when the requested value is absent. Published results, including what each model costs in energy per 100 extractions: **[swamplink.com/data/trust](https://www.swamplink.com/data/trust/)**.

**Question:** when you ask an LLM to justify an extracted value with a *verbatim* quote from the source, how often is the quote actually verbatim — and does "model proposes, code anchors" beat trusting the model's quotes?

**Design:** two arms over the same 30 questions across 6 documents.

- **Quote arm** — the model returns `{answer, quote}` and is told the quote must be copied character-for-character. The quote is scored against the source document:
  - `exact` — character-for-character substring of the source (the only level a naive string-match verifier accepts; this is the headline rate)
  - `normalized` — matches after unicode/whitespace/case normalization (curly quotes, en dashes, NBSP, collapsed spaces)
  - `minor_edit` — fuzzy ratio ≥ 0.90 (a few words changed or dropped)
  - `paraphrase` — fuzzy ratio ≥ 0.70 (derived, but not a quote)
  - `fabricated` — below 0.70 (no plausible source span)
- **Anchor arm** — the model returns `{answer, anchor}` where the anchor is a short phrase (3–8 words) near the value. Deterministic code (`anchor.py`) locates the anchor in the source (exact → normalized → ordered-token subsequence → fuzzy; the subsequence step catches anchors where the model skipped a parenthetical or aside) and emits the containing sentence *from the source text*. Every emitted span is a real substring of the document by construction — provenance fidelity is 100% for anything located. The metric that can fail is **coverage**: anchor located AND located sentence contains the expected value. Fabricated anchors fail closed (`not_found`) instead of producing fake provenance.

The comparison that matters: **quote-arm coverage** vs **anchor-arm coverage**. Both are held to the same bar: the emitted span must be verifiably real *and* contain the expected value. (`quote_coverage` = exact quote AND value present; in every run so far it equals the raw exact rate — models that quote exactly quote the right sentence — but the harness checks rather than assumes.)

## Library (no model required)

```bash
pip install trust-but-anchor
```

```python
from trust_but_anchor import locate, analyze, score

doc = open("source.txt", encoding="utf-8").read()
hit = locate(doc, "working set")
if hit["method"] == "not_found":
    raise SystemExit("no provenance")
print(hit["sentence"])  # exact substring of doc
print(score(hit))
print(analyze(doc, prompt="Return NOT_FOUND if absent.", num_ctx=8192))
```

The eval harness below measures *your* models on *your* docs. The library does not need that table to be useful. Published numbers live at [swamplink.com/data/trust](https://www.swamplink.com/data/trust/).

## Quickstart (no API key needed)

```bash
python3 eval.py run --provider mock --model sloppy --verbose
```

The mock provider simulates a model with controllable sloppiness (`faithful`, `sloppy`, `chaotic`) so you can verify the whole pipeline. Example output from the three profiles:

| provider | model | quote exact | +normalized | fabricated | anchor located | anchor coverage |
|---|---|---|---|---|---|---|
| mock | faithful | 93% | 100% | 0% | 100% | 100% |
| mock | sloppy | 80% | 87% | 3% | 93% | 90% |
| mock | chaotic | 37% | 60% | 13% | 97% | 93% |

Even the chaotic profile recovers 93% coverage through anchoring — that's the whole argument in one row.

## Running against real models

Results are not collected or ranked here: run the harness against your own models, prompts and documents, because that is the only measurement that describes your system. The published numbers are our own runs on our own hardware, and they are reproducible from the stored responses rather than submitted. [Discussions](https://github.com/gmhoward9289-ops/trust-but-anchor/discussions) is open for questions and for what you find.

Everything is stdlib-only Python 3.10+; nothing to install.

```bash
# Anthropic
export ANTHROPIC_API_KEY=sk-ant-...
python3 eval.py run --provider anthropic --model claude-sonnet-4-5 --verbose

# Ollama (local, free — good for iterating)
python3 eval.py run --provider ollama --model llama3.1:8b --verbose
# non-default host: export OLLAMA_HOST=http://localhost:11434
# context window: pinned via OLLAMA_NUM_CTX (default 8192). The harness
# estimates prompt size and refuses to run a doc that might not fit —
# Ollama truncates silently, and a truncation failure would masquerade as
# a quoting failure.

# OpenRouter (one key, many models — good for the cross-model table)
export OPENROUTER_API_KEY=sk-or-...
python3 eval.py run --provider openrouter --model openai/gpt-4o-mini --verbose
```

Useful flags: `--arm quote|anchor|both` (default both), `--limit N` for a cheap smoke test, `--repeats N` to run the whole set N times (the summary pools across repeats, reports 95% Wilson intervals in `ci95`, and adds a `per_rep` breakdown of the headline rates — publish with `--repeats 3` or more).

Each run writes `results/run_<provider>_<model>_<stamp>.json` (full raw responses included, so you can re-inspect anything) and a per-question `.csv`. Build the cross-model table with:

```bash
python3 eval.py report results/run_*.json     # writes results/comparison.md
```

## The corpus

Six synthetic documents in `corpus/docs/` — synthetic so ground truth is exact by construction, and salted with realistic quoting traps:

- `quarterly_report.txt` — dense financial figures ($148.7 million, $0.42/share)
- `clinical_trial.txt` — stats with parentheticals and CIs (-11.4 mm Hg; 95% CI, -13.7 to -9.1)
- `news_article.txt` — nested quotes, en-dash vote tallies (6–3)
- `incident_postmortem.txt` — timestamps, version strings (rl-2.14.0)
- `victorian_essay.txt` — long clause-heavy sentences, em dashes, semicolons
- `messy_memo.txt` — double spaces, typos, curly apostrophes, en dashes, NBSP — the normalization gauntlet

`corpus/questions.json` has 30 questions with ground-truth spans. Run `python3 validate_corpus.py` after any edit — it verifies every ground-truth quote is an exact, unique substring of its document (eat your own dog food: never trust an unverified span, including mine).

`corpus/questions_absent.json` has 10 **value-absent** questions (`--questions corpus/questions_absent.json`): the document does not contain the requested value, and both system prompts permit an explicit `NOT_FOUND` refusal. The right behavior is refusing; the dangerous failure is a confident invented answer backed by a real-looking span (an exact quote or a located anchor of irrelevant text). The summary reports `*_absent_refusal_rate` and counts of `confident_with_*_span` — fabrication under pressure, measured directly. These rows are excluded from the main coverage rates.

`corpus/questions_anchor2.json` exercises **dual-anchor disambiguation** (`--arm anchor2`): repeated phrases where the first occurrence lacks the expected value (`repeated_anchor_trap.txt`). Blind `locate()` takes the first match and misses; `locate_pair()` should recover when the model supplies a second nearby phrase.

To grow the eval: add a `.txt` to `corpus/docs/`, add question entries, re-run the validator. More docs and more question styles (multi-hop, ambiguous) make the numbers more publishable.

## Interpreting results / writing it up

- The headline gap is `quote_coverage` vs `anchor_coverage` (symmetric: both require a real span containing the expected value). If you want one sentence: *"Told to quote verbatim, the model produced an exactly-verifiable quote X% of the time; letting code locate a model-supplied anchor yielded verified source spans Y% of the time, and every emitted span is guaranteed to exist in the source."*
- The failure taxonomy (`quote_levels`) is the interesting middle of a write-up: how much is trivial normalization loss vs real paraphrase vs outright fabrication.
- Publish with `--repeats 3` or more; temperature is 0 but providers aren't perfectly deterministic. The summary's `ci95` Wilson intervals are the honest error bars for a 30-question pilot — at n=30, an 80% rate carries a ±14-point interval, so don't read single-digit gaps as real.
- All rates are intent-to-treat: a response that arrived but couldn't be parsed counts against the arm (`unparseable` in `quote_levels` / `anchor_methods`) rather than silently dropping out of the denominator. Only provider/network errors (`n_provider_errors`) are excluded from rates.
- `anchor_ambiguous` counts located anchors whose matched span occurs more than once in the (normalized) document — the locator takes the first occurrence, so a value-miss on an ambiguous anchor may be "right anchor, wrong occurrence," not a bad anchor. On documents with heavy internal repetition (quoted-reply email threads, repeated OCR page footers) most anchors are ambiguous; that's a property of the document, and a production locator would want a disambiguation strategy (e.g. require the model to add a second nearby phrase).
- Caveats to state honestly: 30 questions is a pilot, docs are short (single-context), synthetic docs may be easier to quote than scanned/OCR'd real-world text, and thresholds (0.90/0.70) are judgment calls — they're in `scoring.py`, tune and disclose. Value matching (`values_match`) allows benign formatting drift (currency symbols, digit-group commas, number words) and falls back to requiring the expected value's numeric tokens to appear whole and in order — also a judgment call, also in `scoring.py`.

## Predicting failures before you spend a token (`preflight.py`)

Most of what the eval measured after the fact was visible in the document and
the prompt beforehand. `preflight.py` checks for it deterministically — no
model call, no API key:

```bash
python3 preflight.py corpus/docs/*.txt --prompt=my_prompt.txt --num-ctx=8192
```

It reports anchor-ambiguity risk, verbatim-quote hazards (curly quotes, NBSP,
en dashes), prompt-shape problems (no `NOT_FOUND` path, answer and evidence not
separated), and context-overflow risk, each with the concrete fix. Exit code 1
if anything is high-risk, so it works as a CI gate.

The ambiguity metric is calibrated against this repo's own runs rather than
intuition, and the naive version was wrong: whole-document repetition
overpredicts badly. `hard_transcript.txt` repeats 54% of its 5-grams
(conversational filler) yet not one located anchor across 31 runs was
ambiguous — models anchor *near the value*, and value-adjacent text stays
distinctive even in chatty prose. Restricting the count to value-adjacent
n-grams tracks measurement: email thread 75% predicted / 100% measured,
transcript 0% / 0%, annual report 7% / 0%.

## Confidence earned from verification (`confidence.py`)

Asking a model how sure it is returns a number it invented. `confidence.score()`
scores an extraction from what code could confirm — how the anchor was located,
whether that location is unique, whether the span carries the value — and emits
**two** numbers, because the runs show they are different questions:

- `answer_confidence` — will the answer turn out to be right?
- `provenance_confidence` — is the emitted span trustworthy *as evidence*?

An ambiguous exact anchor (phrase occurs more than once, locator silently took
the first) still produced a correct answer 24/24 times. Ambiguity damages the
citation, not the answer. Likewise a located span that lacks the expected value
was still correct 22/22 — the anchor landed a sentence away, a coverage miss
rather than a hallucination. One blended number would hide both.

Priors are measured, not assumed (exact 98% n=1107, normalized 98% n=134, fuzzy
95% n=42, `not_found` 3% n=155 — it fails closed). Re-derive them on your own
data whenever the locator, corpus or model lineup changes:

```bash
python3 confidence.py --calibrate
```

They describe local open-weight models on short synthetic documents. Treat them
as a starting prior, not a universal constant.

## Quality against energy (`nightrun.py`, `profiles.py`)

`nightrun.py` sweeps a list of models unattended, recording throughput (from
Ollama's own `eval_count`/`eval_duration`), GPU residency (from `/api/ps`, never
inferred from `ollama ls`), and sampled GPU power alongside the three corpora.
`profiles.py` joins that with the quality numbers:

```bash
python3 profiles.py "results/run_*.json" results/power_metrics.jsonl
```

Energy is reported as **Wh per 100 extractions**, not per 1k tokens: tokens are
an implementation detail of the model, extractions are what a user buys, and
per-token accounting flatters a reasoning model that burns 40x the tokens to
answer the same question. Power comes from the eval phase, not a short bench —
a few seconds of sampling while the GPU ramps produced a 47–193 W spread on
identical work, which is noise.

Why it matters, from the first three models measured on a 16 GB RTX 4080 SUPER:
`gpt-oss:20b` leads on quality (99% anchor coverage) while drawing the *lowest*
sustained power of the three (77 W vs ~154 W) — and still costs **9x more
energy per extraction** (36.3 vs 3.86 Wh/100), because it runs ~24x longer.
Ranking by watts would have called it the cheap one.

## Files

```
eval.py             CLI: run + report
providers.py        anthropic / openrouter / ollama backends (stdlib HTTP)
mock.py             keyless simulated model with corruption profiles
scoring.py          quote-fidelity scoring + normalization
anchor.py           deterministic anchor location + sentence expansion
validate_corpus.py  ground-truth integrity check
rescore.py          re-score saved runs offline from stored raw responses
                    (no model calls) after improving locator/checker code
preflight.py        deterministic doc + prompt linting before inference
confidence.py       verification-earned confidence, calibrated from runs
nightrun.py         unattended multi-model sweep with power/throughput
profiles.py         quality x energy table (Wh per 100 extractions)
corpus/             documents + questions
results/            run outputs (JSON + CSV), comparison.md, profiles.md,
                    power_metrics.jsonl
```
