DocGuard reads what your documentation claims about your code, reads the code, and reports every claim the two no longer agree on — saying which it is sure about, and which it is not.
The faster AI writes code, the faster everyone loses the map of what was built. Every agent session starts from zero and re-derives how the system works. The docs are the map. Nothing checks whether the map still matches the territory.
4. [MED] Traceability DATA-MODEL.md — exists but no matching source code found (unlinked doc) → Doc lives in canonical/ but isn't in the manifest — guard skips it, drift accumulates silently.
npx docguard-cli demo returns seventeen of these in half a second, against a project whose
documentation looks complete. Each names the file, the claim, and what trusting it would cost.
Fig. 1 — The same documentation, under two different questions. DocGuard reports both, and never lets one stand in for the other.
The doc says it; the code agrees. Believe it.
The doc says it; the code says otherwise. DocGuard names the fix.
Something changed near this claim. The judgement is yours — the tool says so, rather than guessing.
A document can be beautifully formatted and completely wrong. A green check is worth exactly as much as the question it answered.
A canonical document holds two kinds of sentence. Some are derived from code — endpoints, entities, environment variables. Some are human reasoning — why the queue is idempotent, which trade-off was rejected, the gotcha nobody would guess. A tool that regenerates the whole file destroys the second kind every time it refreshes the first.
DocGuard keeps both in one readable markdown file, separated by HTML-comment markers. It owns the bytes
inside a source=code block and may rewrite them when the code changes. It never touches anything else.
<!-- docguard:section id=api-endpoints source=code --> | GET | /api/orders/:id | orders.controller.ts | | POST | /api/orders | orders.controller.ts | <!-- /docguard:section --> The orders service is idempotent on client-supplied keys because the mobile client retries on any 5xx; see ADR-007 for the rejected alternative. ← never regenerated
Fig. 2 — One file, two owners. The tool refreshes the blue bands and validates the grey one; it never has permission to do the reverse.
| Direction | Verb | What it does |
|---|---|---|
| code → docs | generate | Reverse-engineer a canonical memory from any codebase. Forty-six scanners build the code-truth skeleton (routes, schemas, screens, env vars, IaC); an agent writes the prose around it. |
| code ↔ docs | sync | When code changes, refresh only the affected source=code sections. Mechanical where deterministic; a structured agent task where not. |
| docs ↔ code | guard | Validate that the memory still matches reality: deleted endpoints still documented, routes never documented, env vars nobody reads, requirement IDs no test traces to. |
Any language: JavaScript/TypeScript and Python get a syntax tree; Rust, Go, Java/Kotlin, Ruby, PHP and C# get a shape-aware fallback, and every finding says which produced it — silence on a fallback file is weak evidence, not a pass.
The tool is allowed to be wrong about the table. It is never allowed to be wrong about your paragraph — so it is not allowed to touch it.
Every run is the same five steps, and none of them calls a language model. The whole pipeline is deterministic, offline, and reproducible — the same tree produces the same findings on every machine, which is what lets a finding be a CI gate rather than an opinion.
Fig. 3 — The pipeline. Step 4 is where DocGuard differs from a linter: a finding is not one number, it is five independent answers.
A finding used to answer one question — how worried should you be? — with one field. That field was doing three jobs: how certain the detector was, whether a human had to look, and whether the number had ever been checked against reality. DocGuard split it into five channels that are independent: a blocking error can still be an escalation, and a high-confidence observation can still be one only you can judge.
| Channel | The question it answers | Values |
|---|---|---|
| severity | Does CI block? | error · warn · info |
| disposition | Who decides — the tool or you? | act DocGuard asserts a defect and names the correction · escalate the judgement is yours |
| confidence | How sure is the detector of its observation? | high · low |
| evidence | Has the reviewed corpus ever measured this code? | measured with n and a Wilson bound · not-measured (a maintainer's prior, nothing more) |
| parserTier | Which analyzer could see this file? | js-ast · py-ast · regex-fallback · fallback-language · mixed |
“13 code commits since ARCHITECTURE.md was reviewed.” The count comes from git log, so it is exact — and it proves only that a review is due, never that the document is wrong. Editing until it stops printing fixes nothing.
“README says 15 validators; the package ships 32.” Both sides are counted from disk, so the tool may rewrite the number — and it stamps the provenance of the count it used, so a later reader can see the fix was about the same subject.
Across fourteen real repositories, 671 findings split 234 the tool would fix and 437 it handed to a human. It says which is which on every line.
Each validator compares one family of documented claims against one source of evidence, grouped here by what they read — which decides how much a finding is worth.
| Family | Reads | Validators | The claim it tests |
|---|---|---|---|
| Code truth · 9 | AST: routes, schemas, env reads, imports | Docs-Sync · Docs-Diff · API-Surface · Schema-Sync · Environment · Architecture · Surface-Sync · Docs-Coverage · Generated-Staleness | “This endpoint / entity / variable / layer rule exists and looks like this.” Both sides read from disk: the findings the tool fixes itself. |
| Structure · 7 | File tree, headings, manifests | Structure (with Doc Sections) · Changelog · Document-Lifecycle · Spec-Registry · Spec-Kit · Metadata-Sync · Doc-Ownership | “Required chapters exist, specs have FR-IDs and phases, versions agree, retired plans left, every path has one owning section.” Completeness, not correctness. |
| References · 5 | Links, symbols, IDs across two revisions | Cross-Reference · Reference-Existence · Traceability · Test-Spec · Path-Scoped-Rules | “This link resolves; this symbol exists; this requirement has a test; a path-scoped rule matches files.” |
| Change & time · 4 | git log, the diff since a ref, code fingerprints | Freshness · Diff-Suspicion · Drift-Comments · Doc-Dependency | “Something changed near this claim.” Exact, but escalated: a review is due. |
| Numbers & self · 3 | Counts on disk, typed JSON, its own package | Metrics-Consistency · Canonical-Sync · Evidence | “The number in this sentence equals the number on disk.” Includes the tool's own README, so its count claims are checked, not remembered. |
| Quality & hygiene · 4 | Prose metrics, secret patterns, TODOs | Doc-Quality · API-Doc-Smells · TODO-Tracking · Security | “Readable, not bloated, not documented in name only; no secret committed; no TODO untracked.” Soft by design. |
A validator with nothing to validate returns N/A with a reason, never a pass — Canonical-Sync outside DocGuard's own repository, Diff-Suspicion without git history. A green line says what it looked at.
Where a published method exists, the detector implements it (page 5). Where none does, the finding's own help text calls the heuristic a heuristic.
First-run conditions on production codebases; every finding hand-classified as true, false, or inert (page 6).
A detector that floods gets precision levers and is re-measured. One that adds nothing is removed — not shipped default-off.
The change-aware detectors are not inventions. Each implements a published result, and each needed a precision lever the paper never had to worry about before it could run unattended in someone else's CI.
| Detector | Published basis | What was taken | What had to be added before default-on |
|---|---|---|---|
| Diff-Suspicion | Outdated-comment detection, arXiv 2010.01625 | A deterministic rule — prose whose tokens overlap a deleted span of the diff is suspect — reached F1 74.7, beating every post-hoc neural model. Applied to docs instead of comments. | Two signals must both hold (the doc references the changed file and shares its removed symbols); a generic-token filter; a per-doc cap. See Fig. 4. |
| Reference-Existence | Two-revision symbol check, arXiv 2212.01479 | A backticked symbol present when the doc was last updated but gone at HEAD is outdated; field-tested at roughly 50% maintainer acceptance. | CLI flags excluded and the two-revision gate made mandatory — the two documented false-positive modes of the source method. |
| API-Doc-Smells | API-documentation smell taxonomy, 1,000-unit benchmark | The two smells with strong deterministic detectors: Bloated (F1 0.90) and Lazy (F1 0.95), keyed on documentation length per signature-headed section. | The three semantic smells that need a BERT model were left out on purpose and routed to staged agent judgement instead. |
| Finding labels | TRACE — calibrated explainability, IEEE TMLCN 2026 | The HIGH / MEDIUM / LOW vocabulary, used as deterministic strata (a validator's check pass-ratio). | An explicit statement that the strata were never calibrated against outcomes — precision is the quantity DocGuard actually measures (page 6). |
| Doc generation | AITPG — multi-agent debate + RAG, IEEE TSE 2026 | The staged prose pattern behind generate and diagnose: the CLI structures, the agent writes, the validators check. | Nothing generated is trusted: every source=code section is re-validated on the next run. |
Evaluating AGENTS.md (arXiv 2602.11988, revised June 2026) found that repository context files did not generally improve agent task success in its evaluated settings, and increased inference cost. It also found agents generally followed the instructions they were given.
DocGuard's README cites this study in its own “Why” section and says plainly that the results
do not establish DocGuard's effectiveness. They motivate two things the tool now does: keep
context concise (retired plans leave the tree; memory packs task-specific context, not everything),
and measure outcomes rather than assume them.
Fig. 4 — Why a paper's F1 is the start, not the end. The same rule, three levers later.
The recall-maximizing variants of these detectors were built, tested, and rejected: a false-positive flood destroys trust faster than a miss does.
Most tools publish an accuracy figure. DocGuard publishes a frozen benchmark, the cases behind it, the confidence interval, the caveat, and a loader that refuses to quote the number if any of those have been edited by hand.
Fig. 5 — The frozen corpus, benchmarks/baseline.json. Every dot is a retained case the loader recomputes from.
Benchmark precision on a deliberately balanced corpus of 12 defect and 12 clean-control cases … This is DocGuard's precision on labelled cases, not the probability that a finding in your repository is real; quote every ratio with its n and Wilson 95% bound.
DocGuard does not publish a calibration document. The real base rate of a stale claim is nowhere near the corpus's 50/50 split, so a probability read off this corpus would mislead — and the README says so in the same paragraph that cites the number.
The baseline is an envelope of retained cases plus derived ratios. The loader recomputes every ratio from the cases and refuses an envelope whose numbers, caveat, or measure have been hand-edited. A test asserts each rejection path; TestGuard probes that the guard cannot be silently removed.
The corpus could only absorb a false positive that had already been repaired — so a disagreement the maintainers reviewed and declined to act on left no trace, and “precision 1.0” described the contribution pipeline rather than the detectors. An adjudication row now records the disagreement without corrupting the measurement.
docguard feedback used to sample only findings the confidence label already doubted — the
channel that would validate the label sampled nothing it could learn from. Selection is now “unmeasured code, or low
confidence”, and the command states what it left out.
“Precision 1.000” is true, and it is a claim about twelve labelled cases. The tool is built to stop that sentence from ever being shortened.
Ninety-three releases in under seven months. The cadence is not the point; what fed it is. Each report from a real working session — most of them written by an AI agent that hit friction on a real tree — was reproduced as a regression test with a control: the pre-fix behaviour, asserted, so a future change cannot silently re-break it.
Fig. 6 — Blue above the line: capability. Clay below: something real broke, and became a test.
DocGuard guards its own repository on every push: guard, score, diff, badge,
and the packed npm tarball run without installed dependencies. Its README count claims are machine-governed. Its own claims are
fault-probed by its sibling, TestGuard. Three things that discipline found:
The semantic-claim extractor ignored .docguardignore; an excluded audit contributed 28 of 39 “unverified claims”. Fixed; the count fell to 12 — and one of those 12 was real drift, a stale suite-runtime claim in TEST-SPEC.md.
REF002 flagged DocGuard's own source: doc comments used realistic ADR-012 examples and fixtures cited ADRs. Non-product scoping and digit-free placeholders; shipped with zero self-findings.
The README said “33 tests” from March 15 through all 92 releases while the suite grew past 2,000 — and every self-guard was green. Metrics-Consistency had computed the test count since v0.8.2 and never compared it to anything. A human found it, not the tool.
declared test cases (static, AST) 1,721 reported by node:test 2,085 // loops generate cases a static count cannot see — so the static number is a FLOOR. MET003 flags a documented count below the floor (stale for certain) and passes anything above it. It carries no mechanical fix: writing the floor would replace a stale number with a wrong one.
A green self-guard proves the checks that exist passed, not that the right checks exist. The honest version of this story — a half-built check, a human reader, a lower-bound design that refuses to guess — is the loop this page is about, and it is why the tool's limitations page comes next.
Through 0.42, DocGuard asked one question: does the documentation still match the code? Version 0.43 adds the questions around it — which section a change should have touched, whether a feature went through its spec first, what an agent needs to read, and whether the tool got slower or noisier. Thirty-eight specs (014–051) were written, planned and tasked before their code; thirty-four carry TestGuard claims probed with planted faults (044 and 048 after this release).
| Area | What changed | Why it matters |
|---|---|---|
| Docs know their code | A section declares the code it describes (covers=); a reviewed fingerprint lock names the section to re-read when that code changes. Module and entity diagrams are drawn from code. | Freshness said a review was due somewhere. This says which paragraph, and why. |
| Spec first, enforced | A spec-first gate classifies each change to governed paths as covered, exempt or uncovered. Completions record the reviewed revision and survive squash merges. As-built specs describe older code. | “We follow Spec Kit” becomes a checked property of the history, not a habit. |
| Agents read less, better | MCP doc tools answer which docs describe this file and return one bounded section. Path-scoped rules report which instruction files each harness loads for a path. | The AGENTS.md study on page 5 measured cost; these cut what an agent must read to find the rule that applies. |
| The tool measured | Budgets run base and head side by side on every PR. An install older than 14 days tells the agent so — from its own CHANGELOG, no network call — and names docguard upgrade. | A guard that slows down or grows noisier erodes trust the same way drift does. |
| Reads real projects | Express, Next.js, FastAPI, Flask, Django, Go, Spring and Rails routes at their served path; read-only commands write nothing. | Defects found on ten real repositories, fixed as specs 040–050. |
The compact guard response was meant to lose nothing. A reconstruction test reads every field back and caught the first draft dropping validator messages.
The symbol map ships opt-in. It becomes a default only if a frozen 54-trial benchmark shows no regression and a benefit on navigation tasks. That run has not happened yet.
Every feature in this release started as a written spec, and nearly every one ended as a probed claim.
.docguard-evidence.json can verify selected statements against typed local evidence; every other statement stays marked unverified, and the count of unverified claims is printed on the badge.guard says whether it is true, and the two are printed apart on purpose.| Alongside | What that does | What DocGuard adds |
|---|---|---|
| GitHub Spec Kit | Generates feature specs, plans, tasks — spec → code, once. | Governance after generation: drift tracking, lifecycle (retire what shipped), spec-quality validation. DocGuard is an official Spec Kit community extension. |
| AGENTS.md | One instruction file: build, test, style. | AGENTS.md is one of the required files. DocGuard validates it, packs task-specific context from the canonical docs, and keeps every agent on the same map. |
| Kiro · Cursor rules | IDE-bound spec or rule files. | Portable across any IDE and any agent; enforced in CI rather than read by one editor. |
| TestGuard | Breaks a promise on purpose and asks whether a test noticed. | The other half of the same question: whether the documentation of that promise is still true. They share an evidence format, and TestGuard fault-probes DocGuard's own claims in CI. |
npm and PyPI packages with provenance attestation, Homebrew, a GitHub Action, a Docker image, an MCP server and Claude Desktop bundle, a Spec Kit extension. No telemetry, no network at validation time.
JSON, SARIF and JUnit with all five channels on every finding; a badge that prints the unverified-claim count next to the check count; a benchmark envelope with its cases attached.
When an agent writes both the code and the docs, the failure mode is not “no docs”. It is confident, plausible documentation of a system that no longer exists. That is exactly what a deterministic comparison against the tree can see.
A score tells you the map is complete. DocGuard tells you which parts of the map still match the territory — and which parts it cannot tell.