Documentation weaknesses and how the refinement campaign addresses them

Decision-support for model-selection-and-global-refine §5. All before/after pairs below are real output from the 25-name stratified pilot (2026-07-17): "before" is the DocsRevision snapshot taken at campaign mark time (the pre-campaign accepted text), "after" is the current accepted text that passed the blind-pair review quorum. Pilot metrics: §5 pilot landed.

§1 — What is wrong with the current docs

About 1,250 accepted entries predate the strict-normative documentation policy. Their docs were written to be helpful narratives; the catalog needs semantic contracts. Four weakness classes, with catalog-wide prevalence (2,288 accepted names currently match at least one defect predicate):

WeaknessPrevalenceWhat it looks likeWhy it is a defect
Typical values1,244 names "Typical single-turn cross-sectional areas … range from $10^{-5}$ to $10^{-3}$ m²"; "representative values are typically of order $10^{-7}$–$10^{-6}$ s" Magnitudes are device-dependent and rot: numbers quoted for one machine class mislead on another (the same doc claimed ±2 m for "compact devices" and ±15 m for "large devices"). A definition must hold for every facility writing the quantity; ranges belong in facility documentation, not the semantic contract. Stale ranges also leak into downstream validation heuristics.
Estimator recipes770 names "In practice, $\tau_A$ is computed from an equilibrium magnetic-axis position…"; "it is obtained from engineering design geometry or from alignment surveys using laser trackers, photogrammetry…" Procedure is not meaning. How a facility estimates a quantity varies by machine, diagnostic, and era; baking one workflow into the definition makes conforming data from a different workflow look non-conforming. The name defines what the quantity is — any estimator that produces that quantity is valid.
Procedural padding671 names "It should be noted…", "Note that…", "In practice…", "For example, …" Filler dilutes normative content. Every non-definitional sentence is a sentence a reviewer, a code generator, or an LLM seat must decide whether to trust — padding raises reading cost without adding a single semantic constraint.
Deterministic audit findings (docs-axis)~135 names latex_def_check: symbol $\phi$ (or $R_0$, $(R,\phi,Z)$) used in a formula with no definition sentence; plus spelling and length-cap findings An undefined symbol makes a formula ambiguous — the reader cannot ground it without guessing conventions. These are mechanically detectable and mechanically verifiable once fixed.

Out of docs scope (important): the largest audit class, decomposition_audit (2,255 accepted names), flags closed-vocabulary tokens absorbed into the name body (e.g. ion in ion_current_density should live in its subject segment). That is a name-axis finding: docs campaigns never touch name identity, so refreshing documentation cannot and should not clear it. 2,125 of the 2,288 matched names carry only name-axis findings — see §4 for what this does to the scale-up choice.

Catalog prevalence (accepted names, 2026-07-17) typical values (1,244)docs-fixable estimator recipes (770)docs-fixable procedural padding (671)docs-fixable docs-axis audit (~139)latex 79 · length 52 · spelling 8 decomposition_audit (2,255)name-axis — NOT docs Pilot clearance (25 names, one gated batch) banned-prose findings: 13 names before → 0 after (grep re-run on refreshed text) · deterministic audit NOT yet re-stamped post-refresh — pre-launch fix queued acceptance 24/25 (96%) · zero prose reintroduction · zero name drift · mean doc length 1,873 → 1,357 chars (−28%)
Prose defects dominate the docs-fixable backlog; the decomposition class is a name-structure problem the docs campaign deliberately does not touch.

§2 — Before / after evidence from the pilot

Example 1 — alfven_time (all three prose classes; 2,458 → 978 chars)

Before (pre-policy)

…Using the magnetic-axis major radius as the characteristic length…, the Alfvén time is defined by: $$\tau_A(\rho_f) = R_0/v_A(\rho_f) = R_0\sqrt{\mu_0\,m_i\,n_i(\rho_f)}/B_0$$ …

In practice, $\tau_A$ is computed from an equilibrium magnetic-axis position, an on-axis toroidal-field magnitude, and ion density and composition at the flux surface of interest. For multi-ion plasmas, the product $m_i n_i$ is commonly replaced by the total ion mass density…

For tokamak parameters $R_0 \sim 1$–$6$ m, $B_0 \sim 1$–$8$ T, and deuterium ion density $n_i \sim 10^{19}$–$10^{20}$ m⁻³, representative values are typically of order $10^{-7}$–$10^{-6}$ s. The timescale increases with major radius…

After (campaign-refined, review score ≥ 0.85)

The Alfvén time is the shear-Alfvén transit time across the magnetic-axis major-radius scale. It is a nonnegative reference timescale for ideal-magnetohydrodynamic dynamics, distinct from the oscillation period or damping time of a particular mode.

$$\tau_A = \frac{R_{\rm axis}}{v_A} = \frac{R_{\rm axis}\sqrt{\mu_0\rho_i}}{|B_{\phi,\rm axis}|}$$

where $R_{\rm axis}$ is the magnetic-axis major radius, $v_A$ is the reference shear-Alfvén speed, $\mu_0$ is the vacuum permeability, $\rho_i$ is the ion mass density at the specified reference location, and $B_{\phi,\rm axis}$ is the toroidal magnetic-field component on the magnetic axis. The magnitude makes the timescale independent of magnetic-field orientation. …distinct from the local complex Alfvén frequency (cross-linked).

The recipe paragraph and the tokamak-parameter range are gone; every symbol in the formula is defined in the where-clause; a disambiguation sentence (reference timescale vs mode period) and a cross-link to alfven_frequency were added. Definition-first, facility-neutral, 60% shorter.

Example 2 — area_of_toroidal_magnetic_field_probe (typical values; 1,491 → 1,066 chars)

Before

…The area is determined from the coil's physical dimensions during fabrication or verified by laboratory calibration; time-integrating the measured voltage and dividing by $NA$ recovers the absolute field $B_\phi$…

Typical single-turn cross-sectional areas in tokamak magnetics diagnostics range from $10^{-5}$ to $10^{-3}$ m², depending on probe dimensions and required sensitivity.

After

The per-turn cross-sectional area of a toroidal magnetic-field probe is the geometric area enclosed by one sensing turn and threaded by the local toroidal magnetic-field component. It is a non-negative geometric parameter, distinct from the total effective area of the complete winding.

$$A_{\mathrm{eff}} = N A_{\mathrm{turn}}$$ where $A_{\mathrm{turn}}$ is the area enclosed by one turn, $N$ is the dimensionless number of turns… The induced voltage of the probe is proportional to the time derivative of the flux coupled through this effective area (cross-linked); analogous to the per-turn area of a poloidal probe (cross-linked).

The calibration workflow and the magnitude range are removed; the ambiguity that actually matters for data consumers — per-turn area vs total winding area — is now the second sentence.

Example 3 — x_coordinate_of_sensor_attachment_point (all three classes; 2,081 → 970 chars)

Before

In practice it is obtained from engineering design geometry or from alignment surveys using coordinate-measuring equipment, laser trackers, photogrammetry, or optical metrology of the installed sensor and mount…

Typical values are set by device size… Compact devices often fall within ±2 m, medium-sized devices within ±5 m, and large devices or external diagnostic supports can reach ±15 m.

After

This quantity gives the x-coordinate of a sensor's mechanical attachment point relative to the origin of the right-handed Cartesian $(x, y, z)$ frame. $$x_{\rm att} = \mathbf{r}_{\rm att} \cdot \hat{x}$$ where $\mathbf{r}_{\rm att}$ is the position vector of the attachment point and $\hat{x}$ the unit vector along the positive x-axis. Together with the y and z coordinates (cross-linked), it specifies the attachment location. A displacement sensor may have two mechanical attachment points; other sensor types have one.

Sign convention: positive when the attachment point lies in the positive x direction from the Cartesian origin.

Survey-equipment inventory and per-device-size ranges deleted; replaced by a projection definition, an explicit multiplicity statement (a genuinely useful semantic fact the old text lacked), and the sign convention retained.

Example 4 — current_of_rogowski_coil (audit-only member; 1,100 → 1,186 chars — docs can also grow)

Before

This quantity is the algebraic net conventional electric current threading the surface bounded by a Rogowski-coil winding contour… $$I_{\mathrm{enc}} = \int_S \mathbf{J} \cdot \hat{\mathbf{n}}\,dS$$ …The effective cross-sectional area of the Rogowski coil characterizes magnetic coupling but is not part of the current definition.

After

…The surface orientation is specified by the coil's assigned normal and the associated right-hand rule for the contour… where $S$ is any surface whose boundary is the coil contour… This is an enclosed-current quantity, not the electrical current carried by the sensor winding. In the Rogowski-coil principle, the induced winding voltage is proportional to $dI_{\mathrm{enc}}/dt$; $I_{\mathrm{enc}}$ is the current whose time variation drives that signal.

Not a compression case: the refinement made the surface-independence of the integral explicit and sharpened the sensor-current vs enclosed-current disambiguation. The campaign target is normativity, not brevity — clean-but-thin docs gain substance (two pilot members with no prior documentation got full docs written).

§2b — Before / after from the live rotations (2026-07-20)

Fresh pairs from rotations 1–2 of the stratified campaign (LLM-adjudicated gate). Same convention: "before" is the DocsRevision snapshot at the pre-rotation accepted text, "after" is the current accepted text that passed the claude-sonnet-5 + grok-4.5 blind-pair quorum. Full ledger: rotation checkpoint ledger.

Example 5 — electron_density_at_magnetic_axis (estimator recipe + typical values; 1,347 → 950 chars, score 0.90)

Before

…$n_e^{\mathrm{axis}} = \lim_{\mathbf{x}\to\mathbf{x}_{\mathrm{axis}}} n_e(\mathbf{x})$ … For a smooth flux-surface quantity, this is equivalently the value at $\rho=0$…

In practice, the magnetic-axis location is determined from an equilibrium reconstruction, and the electron density is evaluated there from local measurements or from a fitted radial density representation. The value is commonly constrained by Thomson scattering, interferometry, or reflectometry…

Typical central electron densities in magnetically confined fusion plasmas are often of order 1019–1020 m−3, varying with fueling, confinement regime, density limits, and profile shape.

After

Particle number density of the electron population evaluated at the magnetic axis is the local electron count per physical volume at the innermost point of nested closed poloidal magnetic-flux surfaces. …$n_e^{\mathrm{axis}} = \lim_{\mathbf{x}\to\mathbf{x}_{\mathrm{axis}}} n_e(\mathbf{x})$ where $n_e(\mathbf{x})$ is the electron number-density field, $\mathbf{x}$ is position, and $\mathbf{x}_{\mathrm{axis}}$ is the magnetic-axis position.

This is the electron-specific realization of density_at_magnetic_axis. It counts free electrons directly and is not summed over ionic charge states or combined with ion populations; ion_density_at_magnetic_axis describes the corresponding ion population.

The measurement-recipe paragraph and the typical-magnitude paragraph are gone; in their place a parent cross-link (density_at_magnetic_axis) and a sibling disambiguation (ion_density_at_magnetic_axis). Definition + family structure replace workflow + numbers.

Example 6 — toroidal_neutral_momentum_diffusivity (recipe + typical values and a physics/notation correction; 2,177 → 1,919 chars, score 0.90)

Before

…the component of the toroidal-momentum flux directed across the transport surface is written as $\Gamma^{\mathrm{diff}} = -D\,\nabla_\perp \Pi_\phi^{\mathrm{n}}$ where $\nabla_\perp$ is the spatial gradient taken normal to the transport surface

This coefficient is generally not measured directly. It is computed or prescribed in fluid-neutral and reduced edge-transport closures from neutral thermal speeds, mean free paths, and momentum-exchange rates…

Typical effective values in edge plasma applications are of order 0.01 to a few m² s⁻¹, depending on neutral temperature, collisionality, geometry…

After

…$\Gamma_{\Pi_j,\phi}^{\mathrm{diff}} = -D_{\mathrm{mom},\phi}^{n}\,\frac{1}{R}\frac{\partial \Pi_j^{n}}{\partial \phi}$ … The factor $R^{-1}\partial/\partial\phi$ is the derivative with respect to arclength in the toroidal direction, and the minus sign gives down-gradient diffusion.

The toroidal qualifier identifies the transport direction (the flux component and gradient direction), not the component $j$ of momentum being transported. This coefficient is one directional component of neutral_momentum_diffusivity, complementary to parallel_, poloidal_, and vertical_neutral_momentum_diffusivity; distinct from the resulting toroidal_neutral_momentum_flux.

Beyond stripping the recipe and the value range, the refinement fixed the physics: the vague "gradient normal to the transport surface" ($\nabla_\perp$) became the correct toroidal derivative $R^{-1}\partial/\partial\phi$, and the transport-direction-vs-momentum-component ambiguity is now explicit — with the full sibling family linked. Normativity, not brevity.

Example 7 — flux_surface_averaged_boron_density_at_plasma_boundary (recipe + typical values removed; per-symbol units added; 2,178 → 1,662 chars, score 0.91)

Before

…$\langle n_\mathrm{B} \rangle_\psi = \frac{\oint_{\psi} n_\mathrm{B}(\ell)\, d\ell / B_p(\ell)}{\oint_{\psi} d\ell / B_p(\ell)}$ where $\ell$ is arclength, $B_p$ is the poloidal magnetic-field magnitude, and $n_\mathrm{B}$ is the local total boron density. (no units stated on the symbols)

In practice, the local boron distribution is inferred from vacuum ultraviolet spectroscopy, charge-exchange recombination spectroscopy, soft X-ray measurements, and collisional-radiative modeling…

Typical flux-surface-averaged boron densities near the confined-plasma edge or separatrix are 1016–1018 m−3, varying with wall conditioning, boronization history…

After

…$\langle n_{\mathrm{B}} \rangle_{\psi_b} = \frac{\oint_{\psi=\psi_b} n_{\mathrm{B}}(\ell)\, d\ell / B_p(\ell)}{\oint_{\psi=\psi_b} d\ell / B_p(\ell)}$ where $\psi$ is the poloidal magnetic flux label with unit Wb, $\ell$ is arclength with unit m, $B_p$ is the poloidal magnetic-field magnitude with unit T, and $n_{\mathrm{B}}$ is the local total boron density with unit m−3.

This quantity is the boundary flux-surface average corresponding to the local boron_density_at_plasma_boundary. It differs from volume_averaged_boron_density and line_averaged_boron_density

Recipe and value-range removed and the boundary/separatrix scoping made explicit — but the refinement also annotated a unit on every symbol (highlighted). Useful precision, or redundant with the structured unit field and the linked quantities? See the open policy question below.

§3 — How each mechanism addresses each weakness

WeaknessMechanism that removes itMechanism that keeps it out
Typical valuesRegeneration under the strict-normative docs prompt (state what the quantity is; no magnitudes, no workflows, no filler) by the gpt-5.6-luna docs seat — benchmarked at 0% banned prose vs 15% for the previous seat (§2 role bench), at rubric parity and 23% lower cost.The same grep-audit vocabulary (prose_policy) is used for campaign selection and the post-batch gate: any refreshed doc still matching a banned pattern counts as reintroduction, and one reintroduction halts the campaign (threshold: zero). A halted batch is root-caused in prompts before resuming — churn is structurally impossible to ignore.
Estimator recipes
Procedural padding
Undefined symbols (latex)The docs rubric requires every symbol in every formula to be defined in a where-sentence; blind-pair review scores this dimension explicitly, and rejected docs loop through refine — one cycle sufficed for nearly every pilot name that needed it (scores rose 0.637 → 0.762 → 1.000 on the worst case).Gap found 2026-07-18: the campaign’s post-batch revalidate currently greps banned prose only — the deterministic audit (latex/spelling/length) is NOT re-run on refreshed docs, and lifted quarantines are confirmed valid on the grep alone. A pre-launch fix routes each batch through the LLM-free validation drain so every refreshed doc is re-stamped and genuine defects re-quarantine.
Physics drift risk (the rewrite damaging correct content)Same accept path as all catalog work — no privileged campaign accept; blind-pair quorum unchanged; every prior text snapshotted to a DocsRevision (reversible); zero-name-drift gate (a campaign can never rename); per-name StandardNameChange events for traceability. Pilot: 0 drift, 96% acceptance.

§4 — What this changes about the scale-up decision

The audit-class breakdown reframes the options presented in the pilot archive. The all spec (2,288 names ≈ $850–960) would spend roughly $350 regenerating docs on ~750 names whose only defect is a name-axis decomposition finding — those docs get freshened, but the flagged defect cannot clear, and that name-structure backlog is already owned by the qualifier-order grammar redesign and individual sn edit renames.

SpecNamesBatchesProjected costCoverage
prose,audit:latex,audit:spelling,audit:length (recommended)1,54016≈ $570–650 as-run; ≈ $300–380 after the refine-livelock fixevery docs-fixable defect: all three prose classes + all docs-axis audit findings
prose1,50416≈ $560–630prose only; leaves 36 docs-axis audit names for hand edits
all (as piloted)2,28823≈ $850–960adds ~750 name-axis-only names whose finding a docs pass cannot clear

Under the recommended spec, "done" for §5 becomes: zero accepted names match any banned-prose pattern (machine-audited), zero docs-axis audit findings on accepted names, and decomposition findings explicitly re-scoped to the name-axis workstream (grammar redesign + per-name renames) rather than waived silently.

§5 — Honest limits

Sources: pilot corpus scratchpad/pilot_before_after.json (session artifact), manifests regenerated 2026-07-18, per-check audit counts from the live graph. Campaign engine: imas-codex a1c8ca1e · pilot machinery 0882e417 · cost telemetry 763a5a14.