API reference
Every public name in the groundlens package, with its parameters,
their types, their defaults and the value to use when you are not sure.
Version 3.0.1.
Installation
The core package has no runtime dependencies. Not numpy, not torch. A CI job
fails the build if that ever changes, so pip install groundlens can
never grow into a deep learning install by accident.
pip install groundlens # core only, no dependencies
pip install "groundlens[encoder]" # + the reference sentence encoder
pip install "groundlens[encoder,mcp]" # + the MCP server
pip install "groundlens[dev]" # + pytest, ruff, mypy
| Extra | Pulls in | Needed for |
|---|---|---|
| — | nothing | The numeral channel, the dataclasses, calibrate(), and any encoder you write yourself. |
| encoder | sentence-transformers>=5.7.0 numpy>=2.2.6 | SentenceTransformerEncoder, which is what scores words. Required for any real use of proofread(). |
| mcp | mcp>=2.0.0 | The MCP server. Combine with encoder; the server needs both. |
| dev | pytest, ruff, mypy | Running the test suite and the linters. |
Python 3.10 through 3.14. Linux, macOS and Windows.
Quickstart
from groundlens import proofread, SentenceTransformerEncoder
answer = "The rate is 4.75% payable within 45 days."
sources = [("policy.pdf#p3", "The rate stated is 3.90% and the term is 30 days.")]
marks = proofread(answer, sources, encoder=SentenceTransformerEncoder(), k=2)
print(marks.report())
# 4.75% support 0.00 nearest in policy.pdf#p3: '3.90%'
# 45 support 0.00 nearest in policy.pdf#p3: '30'
Constructing SentenceTransformerEncoder() downloads the model on
first use, roughly 420 MB, and caches it. Build it once and reuse it: it is the
expensive object in the library.
Concepts
Four terms recur through this reference and are worth fixing before the signatures.
| Term | Meaning |
|---|---|
| anchor | One word of the answer, paired with the best support it found anywhere in the sources, and with the source word that supplied it. |
| support | A number from 0.0 to 1.0. For words it is a cosine similarity. For numerals it is exactly 1.0 or exactly 0.0, because a number is equal or it is wrong. |
| floor | The mean support of the k weakest anchors, not the mean over all of them. A sixty-word answer with one wrong digit has a mean of about 0.79 and a floor of 0.00. |
| receipt | The evidence attached to a mark: which word, where it sits, how it was checked, and the nearest thing in the sources. It is what lets a reviewer settle the call without rereading the document. |
Functions
functiongroundlens.proofread(…)
proofread(
answer: str,
context: str | Sequence[str] | Sequence[tuple[str, str]] | Sequence[Evidence],
*,
encoder: Encoder,
k: int = 1,
locale: str = "und",
max_anchors: int = 2048,
) -> Proofread
Score one answer against its sources and return where to look. This is the function the rest of the library exists to serve.
Parameters
| Name | Type / default | Meaning and recommended value |
|---|---|---|
| answer | str required |
The model output to check. Normalised to NFKC internally, so every character offset in the result refers to the normalised string rather than the one you passed in. |
| context | str Sequence[str] Sequence[tuple[str, str]] Sequence[Evidence] required |
The retrieved sources. Four shapes are accepted: a single string, a list of
strings, a list of (id, text) pairs, or a list of
Evidence.
Use (id, text) pairs. Without an id a finding cannot say which source backed a word, and that is the question a reviewer actually asks. Bare strings are auto-assigned ctx-0, ctx-1 and so on, which tells the reader nothing.
|
| encoder | Encoder required, keyword-only |
Any object satisfying the Encoder protocol.
Use SentenceTransformerEncoder() unless you already run your own retrieval encoder, in which case passing that one makes the check agree with your retrieval. Construct it once and reuse it.
|
| k | int = 1 |
How many of the weakest anchors the floor averages, and how many marks come
back in weakest. k=0 selects
adaptive_k().
Leave it at 1 for review workflows: one word is the only output a person can act on without a second explanation. Use 2 to 4 when you want a short list to skim. Use 0 to reproduce the published benchmark.
|
| locale | str = "und" |
How this corpus writes numbers. See locales.
Set it when you know the corpus. Under "und" an ambiguous numeral such as 1.234 keeps every valid reading, so it matches either 1234 or 1.234 and a real mismatch can slip through. Declaring "es" or "en" removes that slack.
|
| max_anchors | int = 2048 |
Refuse answers with more scoring words than this, rather than quietly taking a very long time. Leave it alone. If you hit the limit, proofread the answer in sections rather than raising it: a floor computed over thousands of words stops pointing anywhere useful. |
Returns
A Proofread. There is no verdict in it and no threshold on it.
Raises
TypeErrorif acontextitem is not a string, an(id, text)pair or anEvidence.ValueErrorif the answer exceedsmax_anchorsscoring words, or iflocaleis not a known name.RuntimeErrorif a word aligns to no encoder token, which means the encoder adapter is broken. It refuses rather than continues, because a silently dropped word can only push the floor up and make a truncated answer look better grounded than it is.
Notes
Passing no context at all is not an error. Every word then reads as unsupported,
which is the correct result, and a warning saying so is attached to
Proofread.warnings.
functiongroundlens.adaptive_k(n)
adaptive_k(n: int) -> int
# max(1, min(4, ceil(0.15 * n)))
How many weakest anchors to use for an answer of n scoring words: one
weak word decides a short answer, up to four decide a long one. Selected by
passing k=0 to proofread(). Call it directly only if you
want the number before you score.
| Name | Type / default | Meaning |
|---|---|---|
| n | int required | Number of scoring words, that is, words not skipped as stopwords or punctuation. |
functiongroundlens.calibrate(…)
calibrate(
labelled: Sequence[tuple[Proofread | float, bool]],
*,
target_recall: float = 0.95,
seed: int = 0,
min_labelled: int = 200,
) -> OperatingPoint
Fit a threshold on your own labelled data and hand back what it costs. This library ships no threshold, and this function is the only supported way to get one.
Parameters
| Name | Type / default | Meaning and recommended value |
|---|---|---|
| labelled | Sequence[tuple[ Proofread | float, bool]] required |
(result_or_floor, is_defect) pairs. You may pass the
Proofread objects directly or just their floor values.
Label your own traffic. A threshold fitted on someone else's distribution does not transfer, because the grounded floor moves with answer length, style and domain.
|
| target_recall | float = 0.95 |
The fraction of defects the cut must catch. Must be in (0, 1].
0.95 for a regulated review, where a missed defect is the expensive error. Lower it toward 0.8 only if you have measured that the false-positive rate at 0.95 makes the queue unworkable, and record that you did.
|
| seed | int = 0 |
Seed for the 1000-round bootstrap that produces the confidence interval. Leave it at 0 so the interval is reproducible and two people fitting on the same data get the same numbers. |
| min_labelled | int = 200 |
The refusal floor. Below this many examples the function raises rather than returning a threshold. Do not lower it to deploy. Lower it only while exploring. Below 200 examples a 95%-recall threshold is estimated from a handful of points, and the interval it produces is not worth reporting. |
Returns
An OperatingPoint carrying the threshold
and the measured false-positive rate with a bootstrap 95% interval.
Raises
ValueErroriftarget_recallis outside(0, 1].ValueErrorif fewer thanmin_labelledexamples are supplied.ValueErrorif the data contains no defects, or no clean examples.
Example
from groundlens import calibrate
point = calibrate(labelled, target_recall=0.95)
print(point.threshold, point.fpr, point.fpr_ci95) # read the fpr first
Across five public RAG grounding benchmarks and nine detectors, including two published encoder models, an NLI cross-encoder, an LLM judge and this metric, not one reached a false-positive rate below 0.65 at 95% hallucination recall. Several sat at 1.00 on benchmarks where they rank well by AUROC. A threshold with an unusable cost is worse than no threshold, which is why this function hands you the bill along with the cut.
Result classes
All of these are frozen dataclasses with slots. They are values, not handles: nothing mutates after proofread() returns.
classgroundlens.Proofread
The result of scoring one answer against its sources.
| Attribute | Type | Meaning |
|---|---|---|
| floor | float | Mean support of the k weakest anchors. The headline number, and the one calibrate() fits on. |
| k | int | The k actually used, after adaptive_k() and after clamping to the number of marked words. |
| weakest | tuple[Anchor, ...] | The k anchors that produced the floor, weakest first. Numerals win ties against words at the same support, because a numeral at 0.0 is a proven mismatch while a word at 0.0 is ordinary in faithful paraphrase. |
| anchors | tuple[Anchor, ...] | Every word of the answer in answer order, including skipped ones. Use this to highlight a whole answer rather than only its worst words. |
| n_marked | int | How many words were scored, excluding skipped ones. |
| n_numeral | int | How many of those were numerals, decided by arithmetic. |
| encoder_id | str | Model name and resolved revision sha, for example all-mpnet-base-v2@bd44305…. Record it with any number you publish. |
| sha256 | str | Content hash of the finding. Covers the structure and the numeral supports exactly, and rounds lexical supports to six decimals. Reproducing the hash reproduces the finding. |
| warnings | tuple[str, ...] | Segmentation and input warnings, for example a largely CJK answer, or no context supplied. Empty when there is nothing to say. |
Methods
| Method | Signature | Meaning |
|---|---|---|
| report | report(limit: int | None = None) -> str | The weakest anchors as receipt lines, newline separated. limit truncates to the first n; None prints all k. |
classgroundlens.Anchor
One word of the answer and the best support it found in the sources.
| Attribute | Type | Meaning |
|---|---|---|
| text | str | The word, verbatim from the NFKC-normalised answer. |
| span | tuple[int, int] | Character offsets into the normalised answer. Use these to highlight in a UI. |
| kind | "lexical" | "numeral" | "skipped" | numeral words were decided by arithmetic, lexical by meaning, skipped were excluded from scoring and always carry support 1.0. |
| support | float | 0.0 to 1.0. Exactly 0.0 or 1.0 for numerals. |
| value | str | None | Canonical decimal string. Numerals only; None otherwise. |
| evidence_id | str | None | Which source supplied the best support. This is the document to open. |
| evidence_text | str | None | The source word that supplied it. This is the receipt. |
| evidence_span | tuple[int, int] | None | Character offsets into that source. |
| notes | tuple[str, ...] | Reason codes. See note codes. |
Methods
| Method | Signature | Meaning |
|---|---|---|
| receipt | receipt() -> str | One tab-separated line a person can act on. Falls back to no anchor found when nothing in the sources matched. |
classgroundlens.OperatingPoint
A threshold measured on your own labelled data, with its cost attached. Returned by calibrate().
| Attribute | Type | Meaning |
|---|---|---|
| threshold | float | Flag an answer when its floor is at or below this value. |
| target_recall | float | The recall you asked for. |
| achieved_recall | float | The recall actually reached on your data. Discreteness means it can differ from the target. |
| fpr | float | False-positive rate at this threshold. Read this first. An fpr of 0.65 means two thirds of your correct answers get escalated. |
| fpr_ci95 | tuple[float, float] | Bootstrap 95% interval on the fpr, over 1000 rounds. |
| n | int | Total labelled examples used. |
| n_positive | int | How many of them were defects. |
Input classes
classgroundlens.Evidence
One retrieved source, with an id so findings can point at it. Construct these yourself when you want explicit typing; (id, text) tuples are converted into them internally.
| Attribute | Type | Meaning and recommended value |
|---|---|---|
| id | str | Identifier that appears in every finding drawn from this source.Make it something a reviewer can open, such as policy.pdf#p3 or a chunk id your retrieval already uses. An opaque hash defeats the purpose. |
| text | str | The passage text as retrieved. |
Encoders
protocolgroundlens.Encoder
The one thing Groundlens needs from a model. It is a runtime-checkable
Protocol, so any object with these four members satisfies it and no
inheritance is required. Implement it to score words with your own retrieval
encoder, which makes the check agree with the retrieval that produced the answer.
| Member | Signature | Contract |
|---|---|---|
| id | property -> str | Stable identity including the exact revision, for example all-mpnet-base-v2@bd44305fd6…. A bare model name is not enough: a silent re-upload would change every number you ever published. This string goes into sha256. |
| max_tokens | property -> int | Hard token limit per window, excluding special tokens. |
| token_spans | token_spans(text: str) -> tuple[Span, ...] | Character spans of every token, without embedding. Windowing needs boundaries before it decides where to cut. Guessing a characters-per-token ratio is how text silently falls off the end of a window and stops being marked at all. |
| encode_window | encode_window(text: str) -> WindowEncoding | Embed one window. The caller guarantees it fits within max_tokens. |
classgroundlens.WindowEncoding
What an Encoder returns for one window of text.
| Attribute | Type | Contract |
|---|---|---|
| token_spans | tuple[Span, ...] | Character offsets of each token, relative to the window's own text. |
| word_ids | tuple[int | None, ...] | Which pre-tokenised word each token belongs to. None for special tokens. |
| vectors | Sequence[Sequence[float]] | Shape (n_tokens, dim), L2-normalised, one row per entry in token_spans. Normalisation is required: support is computed as a plain dot product. |
classgroundlens.SentenceTransformerEncoder
SentenceTransformerEncoder(
model: str = "sentence-transformers/all-mpnet-base-v2",
*,
revision: str | None = None,
device: str | None = None,
max_tokens: int | None = None,
)
The reference encoder, behind the [encoder] extra. Importing
groundlens never imports torch; only constructing this class does.
| Name | Type / default | Meaning and recommended value |
|---|---|---|
| model | str = "sentence-transformers/ all-mpnet-base-v2" |
Any sentence-transformers model id.Use the default unless you are matching your own retrieval encoder, which is the one good reason to change it. Numbers from different encoders are not comparable. |
| revision | str | None = None |
Pin the checkpoint explicitly. When omitted, the sha is resolved from the Hub once at construction; if that fails, the id records "unresolved" rather than pretending.Pin it for anything you publish or audit. An unresolved id in a published hash is a visible defect, which is the point of recording it that way. |
| device | str | None = None |
Passed straight to sentence-transformers, for example "cpu", "cuda", "mps".Leave it None and let the library decide. Results are not bit-identical between x86 and Apple Silicon; they reproduce to 1e-6 with stable ordering, and no stronger claim is made. |
| max_tokens | int | None = None |
Overrides the model's own sequence limit. Two tokens are reserved internally for the special tokens.Leave it alone. Raising it past what the model was trained for degrades the embeddings silently. |
Attributes
| Attribute | Type | Meaning |
|---|---|---|
| id | str | {model-basename}@{revision}, for example all-mpnet-base-v2@bd44305…. |
| max_tokens | int | Resolved content-token limit: the model limit minus the two special tokens. |
Raises
ImportErrorat construction if the[encoder]extra is not installed. The message names the install command.
Type aliases
| Name | Definition | Meaning |
|---|---|---|
| Span | tuple[int, int] | Half-open character offsets, (start, end). |
| AnchorKind | Literal["lexical", "numeral", "skipped"] | Which channel decided a word, or that it was excluded. |
Note codes
groundlens.NOTE_CODES is a frozenset of every code that
can appear in Anchor.notes. The set is closed on purpose so you can
assert on it.
| Code | Meaning |
|---|---|
| numeral_ambiguous | The numeral had more than one valid reading under the declared locale, and matched on the best of them. Declaring a locale removes most of these. |
| numeral_unparsed | It looked numeric but could not be parsed, so it fell back to the lexical channel. |
| window_boundary | The word straddles a window edge; its support is the maximum over the windows it appears in. |
| stopword | Excluded from scoring. |
| no_alpha | Punctuation-only token, excluded from scoring. |
| single_digit | A bare one-digit numeral, treated lexically because these are usually enumerators rather than values. |
Locales
The locale argument tells the numeral channel how this corpus writes
numbers. It is read from the argument and never from LC_ALL, so the
same input gives the same result on any machine.
| Value | Decimal | Grouping | Effect |
|---|---|---|---|
| "und" | — | — | The default. Ambiguous numerals keep every valid reading, so 1.234 matches both 1234 and 1.234. Safe when the corpus is mixed, loose when it is not. |
| "en" | . | , | 1.234 is one and a bit; 1,234 is a thousand. |
| "es" | , | . | 1.234 is a thousand; 1,234 is one and a bit. |
| "de" | , | . | As Spanish. |
| "fr" | , | space | Thousands are separated by a space, including the narrow no-break space. |
| "it" | , | . | As Spanish. |
| "pt" | , | . | As Spanish. |
| "nl" | , | . | As Spanish. |
| "ch" | . | ' | Swiss style, 1'234.50. |
An unknown name raises ValueError listing the known ones.
Constants
| Name | Value | Meaning |
|---|---|---|
| groundlens.__version__ | "3.0.1" | Package version. |
| groundlens.NOTE_CODES | frozenset[str] | Every valid note code. See above. |
| groundlens.calibrate. | 200 | Default refusal floor for calibrate(). |
Command line
Installing the package puts a groundlens command on your path. It has
one subcommand. groundlens --help works without the
[encoder] extra installed, because it does not import torch.
groundlens read --answer ANSWER --context CONTEXT [--context CONTEXT ...]
[--k K] [--locale LOCALE] [--model MODEL] [--json]
| Option | Default | Meaning and recommended value |
|---|---|---|
| --answer | required | The answer text, or a path to a file containing it. A path that exists is read; anything else is treated as the literal text. |
| --context | required, repeatable | A source. Three forms: literal text, a path, or id=path.Use id=path, for example --context policy.pdf#p3=policy.txt, so the output names the document. Without an id the sources become ctx-0, ctx-1. |
| --k | 1 | How many weakest anchors. 0 selects adaptive_k(). |
| --locale | und | One of the locale names. |
| --model | all-mpnet-base-v2 | A sentence-transformers model id. |
| --json | off | Machine-readable output: floor, k, counts, encoder id, sha256, warnings, and every anchor with its receipt.Use it for anything automated. The human format is not a stable interface. |
Example
groundlens read \
--answer answer.txt \
--context policy.pdf#p3=policy.txt \
--locale es --k 2
Warnings go to stderr and the report to stdout. The exit code is 0 on
a successful check; the command does not signal a verdict through its exit code,
because there is no verdict.
MCP server
Groundlens ships an MCP server so Claude Desktop, Claude Code, Cursor, VS Code or any other MCP client can run the check inside the conversation. It speaks stdio and runs locally. No text leaves the machine.
pip install "groundlens[encoder,mcp]"
python -m groundlens.mcp
Client configuration, in claude_desktop_config.json or the equivalent mcp.json:
{
"mcpServers": {
"groundlens": {
"command": "python",
"args": ["-m", "groundlens.mcp"]
}
}
}
Use the absolute path to the Python that has Groundlens installed if it is not the
one on your PATH.
toolfind_unsupported_words
One tool. If a second ever looks necessary, the product has stopped being one thing.
| Parameter | Type / default | Meaning |
|---|---|---|
| answer | string required | The model output to check. |
| sources | array of {id, text} required | The passages the answer was drawn from. Ids appear in the findings. |
| k | integer = 4 | How many of the weakest anchors to return. Higher than the library default, because an assistant can summarise a short list. |
| locale | string = "und" | See locales. |
Returns
| Field | Meaning |
|---|---|
| weakest_anchors | List of {word, support, checked_by, closest_in_sources, source_id, notes}. checked_by is "arithmetic" or "meaning". |
| floor | Mean support of the weakest anchors. |
| n_marked | How many words were scored. |
| encoder_id | Model and revision. |
| sha256 | Content hash of the finding. |
| warnings | Input and segmentation warnings. |
| how_to_read_this | A sentence stating that this is not a verdict, returned with every call so the model cannot omit it. |
The encoder is constructed on the first call rather than at startup, so the server
completes the handshake and answers tools/list with no model on disk.
Exceptions
Groundlens defines no exception classes of its own. It raises built-ins, and each one means something specific.
| Exception | Raised by | Cause |
|---|---|---|
| TypeError | proofread() | A context item was not a string, an (id, text) pair or an Evidence. |
| ValueError | proofread() | Answer above max_anchors, or an unknown locale. |
| ValueError | calibrate() | target_recall outside (0, 1], fewer than min_labelled examples, or missing defect or clean examples. |
| RuntimeError | proofread() | A word aligned to no encoder token. The encoder adapter is broken; continuing would silently raise the floor. |
| ImportError | SentenceTransformerEncoder() | The [encoder] extra is not installed. |
| ImportError | groundlens.mcp | The [mcp] extra is not installed. |
Limitations
- Computed values. It cannot check "revenue tripled" against a source saying revenue went from 5M to 15M.
- Attachment. The word channel checks whether a word is supported, not whether it is attached to the right thing. If an answer says "payable in 30 days" about invoice A and the 30 days belong to invoice B elsewhere in the same context, the word is supported and no mark appears.
- Reasoning. It cannot check an inference. That belongs to entailment models.
- Retrieval. It inherits yours. If the passage is wrong, so is the answer's grounding.
- Segmentation. It assumes space-delimited scripts, and warns rather than pretending when the text is largely CJK or Thai.
The numeral channel is exact: Decimal comparison in a fixed arithmetic context, byte-for-byte identical on any machine, proven in CI on ten operating system and Python combinations under a randomised hash seed and a Turkish locale. The lexical channel is a float32 cosine from a pinned encoder revision; it reproduces to 1e-6 across platforms with stable ordering of the weakest anchors, and is not bit-identical between x86 and Apple Silicon.