API reference

Every public name in the groundlens package, with its parameters, their types, their defaults and the value to use when you are not sure. Version 3.0.1.

Installation

The core package has no runtime dependencies. Not numpy, not torch. A CI job fails the build if that ever changes, so pip install groundlens can never grow into a deep learning install by accident.

pip install groundlens                    # core only, no dependencies
pip install "groundlens[encoder]"         # + the reference sentence encoder
pip install "groundlens[encoder,mcp]"     # + the MCP server
pip install "groundlens[dev]"             # + pytest, ruff, mypy
ExtraPulls inNeeded for
—nothingThe numeral channel, the dataclasses, calibrate(), and any encoder you write yourself.
encodersentence-transformers>=5.7.0
numpy>=2.2.6
SentenceTransformerEncoder, which is what scores words. Required for any real use of proofread().
mcpmcp>=2.0.0The MCP server. Combine with encoder; the server needs both.
devpytest, ruff, mypyRunning the test suite and the linters.

Python 3.10 through 3.14. Linux, macOS and Windows.

Quickstart

from groundlens import proofread, SentenceTransformerEncoder

answer  = "The rate is 4.75% payable within 45 days."
sources = [("policy.pdf#p3", "The rate stated is 3.90% and the term is 30 days.")]

marks = proofread(answer, sources, encoder=SentenceTransformerEncoder(), k=2)

print(marks.report())
#  4.75%   support 0.00    nearest in policy.pdf#p3: '3.90%'
#  45      support 0.00    nearest in policy.pdf#p3: '30'

Constructing SentenceTransformerEncoder() downloads the model on first use, roughly 420 MB, and caches it. Build it once and reuse it: it is the expensive object in the library.

Concepts

Four terms recur through this reference and are worth fixing before the signatures.

TermMeaning
anchorOne word of the answer, paired with the best support it found anywhere in the sources, and with the source word that supplied it.
supportA number from 0.0 to 1.0. For words it is a cosine similarity. For numerals it is exactly 1.0 or exactly 0.0, because a number is equal or it is wrong.
floorThe mean support of the k weakest anchors, not the mean over all of them. A sixty-word answer with one wrong digit has a mean of about 0.79 and a floor of 0.00.
receiptThe evidence attached to a mark: which word, where it sits, how it was checked, and the nearest thing in the sources. It is what lets a reviewer settle the call without rereading the document.

Functions

functiongroundlens.proofread(…)

proofread(
    answer: str,
    context: str | Sequence[str] | Sequence[tuple[str, str]] | Sequence[Evidence],
    *,
    encoder: Encoder,
    k: int = 1,
    locale: str = "und",
    max_anchors: int = 2048,
) -> Proofread

Score one answer against its sources and return where to look. This is the function the rest of the library exists to serve.

Parameters

NameType / defaultMeaning and recommended value
answer str
required
The model output to check. Normalised to NFKC internally, so every character offset in the result refers to the normalised string rather than the one you passed in.
context str
Sequence[str]
Sequence[tuple[str, str]]
Sequence[Evidence]

required
The retrieved sources. Four shapes are accepted: a single string, a list of strings, a list of (id, text) pairs, or a list of Evidence. Use (id, text) pairs. Without an id a finding cannot say which source backed a word, and that is the question a reviewer actually asks. Bare strings are auto-assigned ctx-0, ctx-1 and so on, which tells the reader nothing.
encoder Encoder
required, keyword-only
Any object satisfying the Encoder protocol. Use SentenceTransformerEncoder() unless you already run your own retrieval encoder, in which case passing that one makes the check agree with your retrieval. Construct it once and reuse it.
k int
= 1
How many of the weakest anchors the floor averages, and how many marks come back in weakest. k=0 selects adaptive_k(). Leave it at 1 for review workflows: one word is the only output a person can act on without a second explanation. Use 2 to 4 when you want a short list to skim. Use 0 to reproduce the published benchmark.
locale str
= "und"
How this corpus writes numbers. See locales. Set it when you know the corpus. Under "und" an ambiguous numeral such as 1.234 keeps every valid reading, so it matches either 1234 or 1.234 and a real mismatch can slip through. Declaring "es" or "en" removes that slack.
max_anchors int
= 2048
Refuse answers with more scoring words than this, rather than quietly taking a very long time. Leave it alone. If you hit the limit, proofread the answer in sections rather than raising it: a floor computed over thousands of words stops pointing anywhere useful.

Returns

A Proofread. There is no verdict in it and no threshold on it.

Raises

Notes

Passing no context at all is not an error. Every word then reads as unsupported, which is the correct result, and a warning saying so is attached to Proofread.warnings.

functiongroundlens.adaptive_k(n)

adaptive_k(n: int) -> int
# max(1, min(4, ceil(0.15 * n)))

How many weakest anchors to use for an answer of n scoring words: one weak word decides a short answer, up to four decide a long one. Selected by passing k=0 to proofread(). Call it directly only if you want the number before you score.

NameType / defaultMeaning
nint
required
Number of scoring words, that is, words not skipped as stopwords or punctuation.

functiongroundlens.calibrate(…)

calibrate(
    labelled: Sequence[tuple[Proofread | float, bool]],
    *,
    target_recall: float = 0.95,
    seed: int = 0,
    min_labelled: int = 200,
) -> OperatingPoint

Fit a threshold on your own labelled data and hand back what it costs. This library ships no threshold, and this function is the only supported way to get one.

Parameters

NameType / defaultMeaning and recommended value
labelled Sequence[tuple[
  Proofread | float, bool]]

required
(result_or_floor, is_defect) pairs. You may pass the Proofread objects directly or just their floor values. Label your own traffic. A threshold fitted on someone else's distribution does not transfer, because the grounded floor moves with answer length, style and domain.
target_recall float
= 0.95
The fraction of defects the cut must catch. Must be in (0, 1]. 0.95 for a regulated review, where a missed defect is the expensive error. Lower it toward 0.8 only if you have measured that the false-positive rate at 0.95 makes the queue unworkable, and record that you did.
seed int
= 0
Seed for the 1000-round bootstrap that produces the confidence interval. Leave it at 0 so the interval is reproducible and two people fitting on the same data get the same numbers.
min_labelled int
= 200
The refusal floor. Below this many examples the function raises rather than returning a threshold. Do not lower it to deploy. Lower it only while exploring. Below 200 examples a 95%-recall threshold is estimated from a handful of points, and the interval it produces is not worth reporting.

Returns

An OperatingPoint carrying the threshold and the measured false-positive rate with a bootstrap 95% interval.

Raises

Example

from groundlens import calibrate

point = calibrate(labelled, target_recall=0.95)
print(point.threshold, point.fpr, point.fpr_ci95)   # read the fpr first
Read the false-positive rate before deploying

Across five public RAG grounding benchmarks and nine detectors, including two published encoder models, an NLI cross-encoder, an LLM judge and this metric, not one reached a false-positive rate below 0.65 at 95% hallucination recall. Several sat at 1.00 on benchmarks where they rank well by AUROC. A threshold with an unusable cost is worse than no threshold, which is why this function hands you the bill along with the cut.

Result classes

All of these are frozen dataclasses with slots. They are values, not handles: nothing mutates after proofread() returns.

classgroundlens.Proofread

The result of scoring one answer against its sources.

AttributeTypeMeaning
floorfloatMean support of the k weakest anchors. The headline number, and the one calibrate() fits on.
kintThe k actually used, after adaptive_k() and after clamping to the number of marked words.
weakesttuple[Anchor, ...]The k anchors that produced the floor, weakest first. Numerals win ties against words at the same support, because a numeral at 0.0 is a proven mismatch while a word at 0.0 is ordinary in faithful paraphrase.
anchorstuple[Anchor, ...]Every word of the answer in answer order, including skipped ones. Use this to highlight a whole answer rather than only its worst words.
n_markedintHow many words were scored, excluding skipped ones.
n_numeralintHow many of those were numerals, decided by arithmetic.
encoder_idstrModel name and resolved revision sha, for example all-mpnet-base-v2@bd44305…. Record it with any number you publish.
sha256strContent hash of the finding. Covers the structure and the numeral supports exactly, and rounds lexical supports to six decimals. Reproducing the hash reproduces the finding.
warningstuple[str, ...]Segmentation and input warnings, for example a largely CJK answer, or no context supplied. Empty when there is nothing to say.

Methods

MethodSignatureMeaning
reportreport(limit: int | None = None) -> strThe weakest anchors as receipt lines, newline separated. limit truncates to the first n; None prints all k.

classgroundlens.Anchor

One word of the answer and the best support it found in the sources.

AttributeTypeMeaning
textstrThe word, verbatim from the NFKC-normalised answer.
spantuple[int, int]Character offsets into the normalised answer. Use these to highlight in a UI.
kind"lexical" | "numeral"
| "skipped"
numeral words were decided by arithmetic, lexical by meaning, skipped were excluded from scoring and always carry support 1.0.
supportfloat0.0 to 1.0. Exactly 0.0 or 1.0 for numerals.
valuestr | NoneCanonical decimal string. Numerals only; None otherwise.
evidence_idstr | NoneWhich source supplied the best support. This is the document to open.
evidence_textstr | NoneThe source word that supplied it. This is the receipt.
evidence_spantuple[int, int] | NoneCharacter offsets into that source.
notestuple[str, ...]Reason codes. See note codes.

Methods

MethodSignatureMeaning
receiptreceipt() -> strOne tab-separated line a person can act on. Falls back to no anchor found when nothing in the sources matched.

classgroundlens.OperatingPoint

A threshold measured on your own labelled data, with its cost attached. Returned by calibrate().

AttributeTypeMeaning
thresholdfloatFlag an answer when its floor is at or below this value.
target_recallfloatThe recall you asked for.
achieved_recallfloatThe recall actually reached on your data. Discreteness means it can differ from the target.
fprfloatFalse-positive rate at this threshold. Read this first. An fpr of 0.65 means two thirds of your correct answers get escalated.
fpr_ci95tuple[float, float]Bootstrap 95% interval on the fpr, over 1000 rounds.
nintTotal labelled examples used.
n_positiveintHow many of them were defects.

Input classes

classgroundlens.Evidence

One retrieved source, with an id so findings can point at it. Construct these yourself when you want explicit typing; (id, text) tuples are converted into them internally.

AttributeTypeMeaning and recommended value
idstrIdentifier that appears in every finding drawn from this source.Make it something a reviewer can open, such as policy.pdf#p3 or a chunk id your retrieval already uses. An opaque hash defeats the purpose.
textstrThe passage text as retrieved.

Encoders

protocolgroundlens.Encoder

The one thing Groundlens needs from a model. It is a runtime-checkable Protocol, so any object with these four members satisfies it and no inheritance is required. Implement it to score words with your own retrieval encoder, which makes the check agree with the retrieval that produced the answer.

MemberSignatureContract
idproperty -> strStable identity including the exact revision, for example all-mpnet-base-v2@bd44305fd6…. A bare model name is not enough: a silent re-upload would change every number you ever published. This string goes into sha256.
max_tokensproperty -> intHard token limit per window, excluding special tokens.
token_spanstoken_spans(text: str)
  -> tuple[Span, ...]
Character spans of every token, without embedding. Windowing needs boundaries before it decides where to cut. Guessing a characters-per-token ratio is how text silently falls off the end of a window and stops being marked at all.
encode_windowencode_window(text: str)
  -> WindowEncoding
Embed one window. The caller guarantees it fits within max_tokens.

classgroundlens.WindowEncoding

What an Encoder returns for one window of text.

AttributeTypeContract
token_spanstuple[Span, ...]Character offsets of each token, relative to the window's own text.
word_idstuple[int | None, ...]Which pre-tokenised word each token belongs to. None for special tokens.
vectorsSequence[Sequence[float]]Shape (n_tokens, dim), L2-normalised, one row per entry in token_spans. Normalisation is required: support is computed as a plain dot product.

classgroundlens.SentenceTransformerEncoder

SentenceTransformerEncoder(
    model: str = "sentence-transformers/all-mpnet-base-v2",
    *,
    revision: str | None = None,
    device: str | None = None,
    max_tokens: int | None = None,
)

The reference encoder, behind the [encoder] extra. Importing groundlens never imports torch; only constructing this class does.

NameType / defaultMeaning and recommended value
model str
= "sentence-transformers/
  all-mpnet-base-v2"
Any sentence-transformers model id.Use the default unless you are matching your own retrieval encoder, which is the one good reason to change it. Numbers from different encoders are not comparable.
revision str | None
= None
Pin the checkpoint explicitly. When omitted, the sha is resolved from the Hub once at construction; if that fails, the id records "unresolved" rather than pretending.Pin it for anything you publish or audit. An unresolved id in a published hash is a visible defect, which is the point of recording it that way.
device str | None
= None
Passed straight to sentence-transformers, for example "cpu", "cuda", "mps".Leave it None and let the library decide. Results are not bit-identical between x86 and Apple Silicon; they reproduce to 1e-6 with stable ordering, and no stronger claim is made.
max_tokens int | None
= None
Overrides the model's own sequence limit. Two tokens are reserved internally for the special tokens.Leave it alone. Raising it past what the model was trained for degrades the embeddings silently.

Attributes

AttributeTypeMeaning
idstr{model-basename}@{revision}, for example all-mpnet-base-v2@bd44305….
max_tokensintResolved content-token limit: the model limit minus the two special tokens.

Raises

Type aliases

NameDefinitionMeaning
Spantuple[int, int]Half-open character offsets, (start, end).
AnchorKindLiteral["lexical",
  "numeral", "skipped"]
Which channel decided a word, or that it was excluded.

Note codes

groundlens.NOTE_CODES is a frozenset of every code that can appear in Anchor.notes. The set is closed on purpose so you can assert on it.

CodeMeaning
numeral_ambiguousThe numeral had more than one valid reading under the declared locale, and matched on the best of them. Declaring a locale removes most of these.
numeral_unparsedIt looked numeric but could not be parsed, so it fell back to the lexical channel.
window_boundaryThe word straddles a window edge; its support is the maximum over the windows it appears in.
stopwordExcluded from scoring.
no_alphaPunctuation-only token, excluded from scoring.
single_digitA bare one-digit numeral, treated lexically because these are usually enumerators rather than values.

Locales

The locale argument tells the numeral channel how this corpus writes numbers. It is read from the argument and never from LC_ALL, so the same input gives the same result on any machine.

ValueDecimalGroupingEffect
"und"——The default. Ambiguous numerals keep every valid reading, so 1.234 matches both 1234 and 1.234. Safe when the corpus is mixed, loose when it is not.
"en".,1.234 is one and a bit; 1,234 is a thousand.
"es",.1.234 is a thousand; 1,234 is one and a bit.
"de",.As Spanish.
"fr",spaceThousands are separated by a space, including the narrow no-break space.
"it",.As Spanish.
"pt",.As Spanish.
"nl",.As Spanish.
"ch".'Swiss style, 1'234.50.

An unknown name raises ValueError listing the known ones.

Constants

NameValueMeaning
groundlens.__version__"3.0.1"Package version.
groundlens.NOTE_CODESfrozenset[str]Every valid note code. See above.
groundlens.calibrate.MIN_LABELLED200Default refusal floor for calibrate().

Command line

Installing the package puts a groundlens command on your path. It has one subcommand. groundlens --help works without the [encoder] extra installed, because it does not import torch.

groundlens read --answer ANSWER --context CONTEXT [--context CONTEXT ...]
                [--k K] [--locale LOCALE] [--model MODEL] [--json]
OptionDefaultMeaning and recommended value
--answerrequiredThe answer text, or a path to a file containing it. A path that exists is read; anything else is treated as the literal text.
--contextrequired,
repeatable
A source. Three forms: literal text, a path, or id=path.Use id=path, for example --context policy.pdf#p3=policy.txt, so the output names the document. Without an id the sources become ctx-0, ctx-1.
--k1How many weakest anchors. 0 selects adaptive_k().
--localeundOne of the locale names.
--modelall-mpnet-base-v2A sentence-transformers model id.
--jsonoffMachine-readable output: floor, k, counts, encoder id, sha256, warnings, and every anchor with its receipt.Use it for anything automated. The human format is not a stable interface.

Example

groundlens read \
  --answer answer.txt \
  --context policy.pdf#p3=policy.txt \
  --locale es --k 2

Warnings go to stderr and the report to stdout. The exit code is 0 on a successful check; the command does not signal a verdict through its exit code, because there is no verdict.

MCP server

Groundlens ships an MCP server so Claude Desktop, Claude Code, Cursor, VS Code or any other MCP client can run the check inside the conversation. It speaks stdio and runs locally. No text leaves the machine.

pip install "groundlens[encoder,mcp]"
python -m groundlens.mcp

Client configuration, in claude_desktop_config.json or the equivalent mcp.json:

{
  "mcpServers": {
    "groundlens": {
      "command": "python",
      "args": ["-m", "groundlens.mcp"]
    }
  }
}

Use the absolute path to the Python that has Groundlens installed if it is not the one on your PATH.

toolfind_unsupported_words

One tool. If a second ever looks necessary, the product has stopped being one thing.

ParameterType / defaultMeaning
answerstring
required
The model output to check.
sourcesarray of {id, text}
required
The passages the answer was drawn from. Ids appear in the findings.
kinteger
= 4
How many of the weakest anchors to return. Higher than the library default, because an assistant can summarise a short list.
localestring
= "und"
See locales.

Returns

FieldMeaning
weakest_anchorsList of {word, support, checked_by, closest_in_sources, source_id, notes}. checked_by is "arithmetic" or "meaning".
floorMean support of the weakest anchors.
n_markedHow many words were scored.
encoder_idModel and revision.
sha256Content hash of the finding.
warningsInput and segmentation warnings.
how_to_read_thisA sentence stating that this is not a verdict, returned with every call so the model cannot omit it.

The encoder is constructed on the first call rather than at startup, so the server completes the handshake and answers tools/list with no model on disk.

Exceptions

Groundlens defines no exception classes of its own. It raises built-ins, and each one means something specific.

ExceptionRaised byCause
TypeErrorproofread()A context item was not a string, an (id, text) pair or an Evidence.
ValueErrorproofread()Answer above max_anchors, or an unknown locale.
ValueErrorcalibrate()target_recall outside (0, 1], fewer than min_labelled examples, or missing defect or clean examples.
RuntimeErrorproofread()A word aligned to no encoder token. The encoder adapter is broken; continuing would silently raise the floor.
ImportErrorSentenceTransformerEncoder()The [encoder] extra is not installed.
ImportErrorgroundlens.mcpThe [mcp] extra is not installed.

Limitations

Reproducibility

The numeral channel is exact: Decimal comparison in a fixed arithmetic context, byte-for-byte identical on any machine, proven in CI on ten operating system and Python combinations under a randomised hash seed and a Turkish locale. The lexical channel is a float32 cosine from a pinned encoder revision; it reproduces to 1e-6 across platforms with stable ordering of the weakest anchors, and is not bit-identical between x86 and Apple Silicon.