Lines
Saved output line count.
A local-first release command center that turns release notes, docs, validation runs, screenshots, package artifacts, risks, waivers, and go/no-go decisions into one auditable workspace.
| Variant | Output | Lines | Words |
|---|---|---|---|
| [C1] Manual brief | 00-control-manual-release-readiness-workbench.md | 9 | 55 |
| [C2] Matched reusable-template control | 01-control-template-release-readiness-workbench.md | 151 | 1,069 |
| [T] WeaveMark treatment | 02-treatment-promplet-release-readiness-workbench.md | 575 | 4,762 |
Saved output line count.
Saved compiled output word count.
Verbatim source/output material quoted from saved artifacts.
Design a local-first web application that helps prepare a software release. It should collect release tasks, evidence, validation results, docs, examples, risks, and go/no-go decisions in one workspace.
limitations, linked claim, and release impact. - Critical gates should block release unless explicitly waived with rationale, owner, approver, expiry, accepted risk, and revisit trigger. - Validation checks need expected proof, owner, status, last run, failure meaning,
Build **Release Readiness Workbench**, a local-first release command center that turns messy release material into structured gates, evidence, validation, risks, actions, notes, and go/no-go records. This is a directly implementable TypeScript/Next.js/Prisma/SQLite application specification for an AI programming agent. The result is a single local workspace for proving whether a release is ready across WeaveMark public README, docs, examples, generated outputs, study results, release notes, package artifacts, extension builds, CLI behavior, installation checks, single-output validation studies, qualitative evidence, score explanations, browser-facing examples, screenshots, traces, console/runtime findings, issues, PR notes, local TODOs, deferred work, waivers, and known limitations.
@refine programming/foundations/software-spec Mingle refinements into one release workspace; no appendices.
Scores compare [T] WeaveMark treatment against [C2] Matched reusable-template control on the -3..+3 scale.
Primary scores are blind* using hybrid-derived-metrics-and-masked-review: anonymous absolute 1..7 scores were frozen before reveal, then converted to the -3..+3 treatment-control scale. Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.
| Criterion | Blind* score | Evidence |
|---|---|---|
| Authoring leverage | +2 | derived-evidence method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control. |
| Information yield | -2 | derived-evidence method. Blind* absolute scores: 4 for [T] versus 7 for the strongest control. |
| Grounded expressiveness | +2 | masked-source-output review method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control. |
| Input readability | +1 | masked-source review method. Blind* absolute scores: 5 for [T] versus 4 for the strongest control. |
| Output readability | +1 | masked-output review method. Blind* absolute scores: 6 for [T] versus 5 for the strongest control. |
| Constraint integration | +2 | masked-source-output review method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control. |
| Reusable abstraction quality | +1 | masked-source review method. Blind* absolute scores: 6 for [T] versus 5 for the strongest control. |
| Total | +7 | Net contrastive gain/loss. |
A -3..+3 judgment comparing [T] against the strongest listed control for each criterion.
The sum of the seven contrastive criterion scores for one study.
Green means [T] is better, red means worse, amber means similar; intensity follows magnitude.
A strong headline study, with the honest caveat that the template remains denser and more source-efficient on the yield proxy. The qualitative claim should include both sides: WeaveMark improves semantic integration where shown, but the measured failures and caveats are part of the result.