Metadata-Version: 2.4
Name: stillvalid
Version: 0.1.1
Summary: Is this information still valid? A validity layer between retrieval and action — deterministic checks first, learned survival probability where nothing deterministic exists.
License: MIT
Project-URL: Homepage, https://github.com/junniec01-creator/stillvalid
Keywords: freshness,staleness,rag,agents,validity,cache,survival-analysis
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: mcp
Requires-Dist: mcp>=1.0; extra == "mcp"
Provides-Extra: langchain
Requires-Dist: langchain-core>=0.3; extra == "langchain"
Provides-Extra: llamaindex
Requires-Dist: llama-index-core>=0.11; extra == "llamaindex"
Dynamic: license-file

# stillvalid

**Is this information still valid?**

A validity layer between retrieval and action. Your RAG pipeline or agent
retrieved a document — `stillvalid` tells you whether to trust it, verify
it, or refresh it, *before* the LLM acts on stale facts.

> Humans pause when they notice "this doc is from 2023". Agents don't —
> they read stale pricing and generate the quote. A tribunal has already
> held a company liable for its chatbot citing an outdated policy
> (*Moffatt v Air Canada*, 2024). Staleness is becoming an agent-safety
> problem, not a search-quality nit.

```python
from stillvalid import Checker, Doc

sv = Checker(history_db="observations.db")

verdicts = sv.check_many(
    Doc(id=d.id, last_verified_at=d.indexed_at, last_changed_at=d.modified_at)
    for d in retrieved_docs
)
for d, v in zip(retrieved_docs, verdicts):
    if not v.usable:           # VERIFY / STALE / UNKNOWN
        d = refetch(d)         # re-check only the risky evidence
```

Standard library only. Zero dependencies, zero network calls, microseconds
per check. **Experimental (v0.1) — API may change.**

## 30-second demo

```
python examples/demo.py
```

A support agent indexed its knowledge base 45 days ago; the refund policy
changed upstream 10 days ago. Without a validity layer the agent quotes
the old fee *and* an expired promo. With it:

```
document          state         layer        reason
policy/refund     STALE         modified     source was modified 35d after your last verification
promo/summer      STALE         expiry       explicitly expired 5d ago
policy/baggage    VALID         hash         live content hash equals verified hash
fees/schedule     STALE         survival     calibrated 21% probability it is unchanged 45d after verification
guide/visa        LIKELY_VALID  survival     calibrated 94% probability it is unchanged 45d after verification

Agent re-fetches only the risky evidence: refund, promo, fees
  → 2/5 documents served from cache (no re-fetch cost);
    3/5 re-verified — the wrong answer never left the building.
```

Every verdict also carries a one-sentence `summary` built for agents to
relay ("'policy/refund' is STALE — source was modified 35d after your
last verification. Re-fetch the source before quoting it.") — so in MCP
use, the check narrates itself in the conversation.

Note the last two rows: both are "45 days since verification", but the fee
table's learned update habit says *stale* while the visa guide's says
*still fine* — that distinction is the whole point, and no timestamp
cutoff can make it.

## How it decides: a cascade

Cheap, certain judgments first; stop at the first layer that can decide;
abstain honestly when none can.

| # | Layer | Needs | Verdict type |
|---|-------|-------|--------------|
| 1 | Explicit expiry (`expires_at`) | nothing | deterministic |
| 2 | Live content-hash compare | a hash you just fetched | deterministic |
| 3 | Live Last-Modified compare | a timestamp you just fetched | deterministic |
| 6 | **Calibrated survival model** | gate-passed params for this doc | probability, `calibrated=True` |
| 5 | Crude change-rate estimate | ≥ 2 observed changes in local history | probability, `calibrated=False` |
| 7 | `UNKNOWN` | — | honest abstention |

Two design rules worth knowing:

- **If something deterministic exists, probability never runs.** When you
  can simply look the answer up, a model is the wrong tool. The learned
  layers exist for the documents where no lookup can answer — sources
  without feeds, and the question "will it *still* be valid when I act?",
  which no feed can answer.
- **Uncalibrated numbers say so.** Layer 5 is a memoryless estimate that
  exists so the cascade is useful on day one; its verdicts carry
  `calibrated=False`. Layer 6 probabilities come from a survival model
  that must pass a serving gate (discrimination + ablation + calibration
  checks) before a single probability is published. Documents that
  haven't earned a calibrated number get `UNKNOWN`, not a guess.

## URL sources work out of the box

For anything with a URL, the deterministic layers don't need you to wire
up anything — one conditional HEAD request (no body, no re-embedding)
fetches the signals:

```python
from stillvalid.probe import probe, doc_from_probe

r = probe(url, etag=indexed_etag, if_modified_since=indexed_at)
verdict = sv.check(doc_from_probe(doc_id, r,
                                  last_verified_at=indexed_at,
                                  verified_etag=indexed_etag))
```

A `304 Not Modified` is the cheapest deterministic VALID there is; a
changed ETag or newer `Last-Modified` is a deterministic STALE. Probe
failures degrade to UNKNOWN — a network error is not evidence of
staleness. `probe_many([...])` does batches concurrently.

In a 25-site survey (docs, APIs, dynamic pages), **74% exposed an ETag or
`Last-Modified`** — so the deterministic layers carry roughly three out of
four real-world URLs. Two field notes from that survey:

- **Use ETags, not body hashes.** Every dynamic page we measured changed
  its bytes on *every* request (CSRF tokens, analytics ids) while its
  content was identical. Feeding a whole-body hash into `verified_hash`
  would manufacture constant false STALE; the server's own validator does
  not have that problem.
- **Watch `result.age`.** Validators can arrive from a CDN cache — we saw
  one up to 5.3 hours old, and `Cache-Control: no-cache` does not defeat
  it (CDNs ignore it). A cached validator describes the origin as of
  `age` seconds ago, not now. Pass `doc_from_probe(..., max_age_s=…)` to
  drop validators older than you are willing to trust; they then fall
  through to the probabilistic layers or UNKNOWN.

### Probing is locked down by default

URLs usually arrive from a retriever — that is, from *data* — so the
probe treats them as untrusted:

- **http/https only.** Left unguarded, `urllib` happily serves `file://`
  and `ftp://`, which would turn "check this document" into local file
  disclosure.
- **Public addresses only.** Loopback, private ranges, link-local
  (including cloud metadata at `169.254.169.254`), reserved and multicast
  targets are refused.
- **Redirects re-checked at every hop**, so a public URL answering `302`
  with an internal `Location` does not become an escape hatch.
- **No credentials are ever attached** — no cookies, no auth headers.

Internal sources are a legitimate case (a wiki on `10.x` is normal), so
they are opt-in, not silently allowed:

```python
probe(url, allow_private=True)     # I mean it, this host is internal
```

Known limit: the host is resolved once for the check and again by
`urllib` for the connection, so DNS rebinding is not covered — pin your
own resolver if that is in your threat model.

**What gets stored**: the observation log keeps only `doc_id`, a
timestamp, and the content hash you passed. `doc_id` is stored verbatim,
so if you pass raw URLs, anything embedded in them (tokens, credentials,
internal paths) is stored too. Pass an opaque id if that matters.

## Import the past instead of waiting for it

The usual deal with a freshness layer is "install it, come back in a
month". Skip that: most knowledge bases already keep their own change
history, and reading it is a one-liner.

```python
from stillvalid.backfill import from_git, from_csv, from_rows

from_git(sv, r"C:\work\handbook")          # a docs repo: commits are the log
from_csv(sv, "wiki_revisions.csv")         # doc_id, ts[, rev]
from_rows(sv, confluence_revision_rows)    # anything you can iterate
```

Measured on three real repositories here: **3,228 past changes across
1,089 documents, imported in 1.7 seconds.** Before the import, a sample of
those documents judged `UNKNOWN` across the board; afterwards the same
query returned real verdicts with probabilities spread from 0% to 88%.
That is the difference between "ask me again in a month" and "ask me now".

Git is the easiest case, but the general entry point is `from_rows` —
feed it a wiki's revision API, a CMS audit table, a warehouse query of
`updated_at` snapshots. Consecutive identical revisions read as "still the
same", which is exactly the censoring signal the learned layers want.

`summarize(sv)` tells you what the import bought: how many documents can
now be judged on update behavior, and how many have enough events (8+) to
train a calibrated model on.

## It learns as you use it

The LangChain / LlamaIndex wrappers **auto-record** a content-hash
sighting for every document that flows through them (same-hash sightings
are deduped to one per hour; disable with `record=False`). Or log
manually whenever you fetch or reindex:

```python
sv.record(doc_id, content_hash)
```

Day 1: layers 1–3 and 7 carry the load. Within days, layer 5 starts
estimating from your observation log. With enough history (8+ observed
changes per document), you can train the survival model and export
parameters — then "updated 7 days ago" upgrades to "97% likely still
valid", calibrated.

One honesty note on auto-recording: it observes *your index's copy*, so
a change becomes visible at reindex time, not at source-change time.
That is exactly the fidelity the crude layer claims — calibrated
training should prefer source-side timestamps where available.

The trainer (survival analysis over your observation log, with the
serving gate) lives in the parent project; `stillvalid` only needs its
exported JSON (`fpi-qtime-params/1`).

## Cost

Measured on one laptop (Windows, Python 3.12) — order of magnitude, not a
benchmark claim:

| Operation | Cost |
|---|---|
| Verdict, deterministic layers | ~3 µs |
| Verdict, change-rate layer | ~0.5 ms, **constant** in history size |
| Batch of 1,000 documents | ~7 ms |
| `live=True`, 50 local files | ~0.9 ms |
| `live=True`, remote URLs (batched) | **~300 ms**, one round trip |
| Recording a sighting | ~0.4 ms (deduped: ~0.02 ms) |
| Observation log on disk | ~52 bytes/row |

Two things worth planning around:

- **Remote `live=True` costs a network round trip** (~300 ms, flat for a
  batch since probes run concurrently). That is fine for an agent deciding
  whether to act, and probably too slow in a latency-critical search path —
  there, prefer local paths, or refresh signals on a schedule rather than
  per query.
- **The change-rate layer only reads the most recent sightings**
  (`History.RECENT_WINDOW`, 500). Reading full history made judging cost
  grow with it; recent behavior also predicts better, so the window is
  both faster and more honest.

## Verdicts

| State | Meaning | Default action |
|-------|---------|----------------|
| `VALID` | deterministically confirmed current | `use` |
| `LIKELY_VALID` | probability ≥ use-threshold (default 0.90) | `use` |
| `VERIFY` | uncertain — recheck the source | `verify` |
| `STALE` | expired / changed / probably invalid | `refresh` |
| `UNKNOWN` | no layer could judge | `verify` |

Thresholds are yours to set: `Checker(policy=Policy(use_threshold=0.95,
stale_threshold=0.70))`. Every verdict carries the deciding layer, a
human-readable reason, and the probability when one exists — evidence,
not just a verdict.

## Why thresholds on a calibrated probability (and not doc age)?

Because age can rank documents, but it can't answer "is 30 days old fine
*for this document*?" — a HQ address and a stock level are both "30 days
old" and deserve opposite treatment. In our backtests on a public-document
corpus (300k queries, weekly-reindex scenario), threshold decisions on
calibrated probabilities left 2.4–4.8× less stale exposure than an
age-cutoff heuristic at the same re-verification budget.

## Integrations

**LangChain** — judge every retrieved document between retrieval and the
LLM (`pip install "stillvalid[langchain]"`):

```python
# LangChain >= 1.0 moved this retriever into langchain_classic
from langchain_classic.retrievers import ContextualCompressionRetriever
from stillvalid import Checker
from stillvalid.integrations.langchain import StillValidCompressor

retriever = ContextualCompressionRetriever(
    base_compressor=StillValidCompressor(checker=Checker(...), live=True),
    base_retriever=vectorstore.as_retriever(),
)
# each doc gains metadata["stillvalid"] = {state, action, layer, reason, ...}
```

**LlamaIndex** — same idea as a node postprocessor
(`pip install "stillvalid[llamaindex]"`):

```python
from stillvalid.integrations.llamaindex import StillValidPostprocessor

engine = index.as_query_engine(
    node_postprocessors=[StillValidPostprocessor(checker=Checker(...))])
```

Both read common metadata keys out of the box (`indexed_at`,
`last_modified`, `updated_at`, `content_hash`, LlamaIndex's
`last_modified_date`, …) and `stillvalid_*` prefixed keys always win.
`mode="filter"` drops non-usable documents but keeps `UNKNOWN` by
default — absence of evidence is not evidence of staleness
(`drop_unknown=True` if you disagree).

**Pass `live=True` unless you have a reason not to.** Index metadata is a
*past* snapshot by definition — it says when you last looked, never
whether the source has moved since — so without a current signal a fresh
install can only answer UNKNOWN. `live=True` fetches that signal at query
time, cheaply: `os.stat()` for local paths, a batched conditional HEAD for
URLs (never the body). In our end-to-end run over a real
loader → vector store → retriever pipeline, this turned *every* verdict
from UNKNOWN into a real one and flagged exactly the file that had been
edited after indexing — at **2 ms** added latency for local files.
Network sources cost one round trip; tune with `live_timeout`, and set
`allow_private=True` if your wiki is on an internal address.

**MCP server** — let the agent itself ask
(`pip install "stillvalid[mcp]"`, supports mcp 1.x and 2.x):

```jsonc
// Claude Desktop / any MCP client
{"mcpServers": {"stillvalid": {
    "command": "python", "args": ["-m", "stillvalid.mcp_server"],
    "env": {"STILLVALID_HISTORY_DB": "C:/data/observations.db"}}}}
```

Tools: `import_history` (load a corpus's past in seconds — run this
first), `check_validity` (judge one piece of evidence before acting on
it), `check_validity_batch` (shared clock), `record_observation` (log a
sighting so the cascade keeps learning).

In an agent session that looks like: *"point me at your docs repo"* →
`import_history` → *"3,228 past changes imported; 554 of 1,089 documents
can now be judged on their update behavior"* → every later answer is
checked against that.

## Roadmap

- Cross-source verification layer (layer 4 of the design)
- Cold-start priors: borrowing update-behavior from similar sources,
  gated against negative transfer
- `pip install stillvalid` (PyPI release)

## Status & honesty

This is an extraction of a working research pipeline (survival analysis
over change histories, with calibration verification) into a standalone
layer. The calibrated path is real but requires training on your
observation log; everything else works out of the box. If your documents
all have reliable feeds or explicit expiries, you don't need the learned
layers — and this library will happily tell you so by never reaching them.
