WeaveMark qualitative report

Verdant Relay Final Quality Analysis

A browser game about defending a living railway garden from blight by combining tower-defense route pressure, deckbuilder card choices, ecosystem feedback, original assets, and browser validation.

Outputs inspected

VariantOutputLinesWords
[C1] Manual brief00-control-compact-manual-verdant-relay.md1057
[C2] Matched reusable-template control01-control-template-verdant-relay.md139954
[T] WeaveMark treatment02-treatment-promplet-verdant-relay.md6655,567
Metric definitions and scoring legendOpen for exact meanings.

Lines

Saved output line count.

Words

Saved compiled output word count.

Verbatim snippets

Verbatim source/output material quoted from saved artifacts.

[C1] Manual briefOpen output
Design a browser game called Verdant Relay that combines tower defense, deckbuilding, and ecosystem simulation. The player protects a living railway garden from spreading blight by placing defenses, playing cards, and maintaining ecological balance.
[C2] Matched reusable-template controlOpen output
- This is the coherent implementation-ready specification for Verdant Relay. - Treat this specification as the source of truth for a programming agent or human engineer. - Include first-build scope, out-of-scope items, architecture, domain model, durable records, workflows, UI surfaces, automation rules, validation plan,
[T] WeaveMark treatmentOpen output
Build **Verdant Relay**, a playable first-build browser game where the player protects a living railway corridor from blight waves by placing and upgrading ecological defenses, playing deckbuilder cards, and maintaining ecosystem health. The specification is the source of truth for implementation, validation, tuning, assets, and acceptance.
[T] WeaveMark treatmentOpen source seam
@compress "Produce a dense browser-game implementation spec; preserve structure and every hard requirement." @refine programming/foundations/software-spec Mingle mechanics, architecture, assets, validation, tuning, and readability into one browser-game implementation spec; never append generic fragments.

Contrastive gain/loss scores

Scores compare [T] WeaveMark treatment against [C2] Matched reusable-template control on the -3..+3 scale.

Primary scores are blind* using hybrid-derived-metrics-and-masked-review: anonymous absolute 1..7 scores were frozen before reveal, then converted to the -3..+3 treatment-control scale. Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.

CriterionBlind* scoreEvidence
Authoring leverage+2derived-evidence method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control.
Information yield+2derived-evidence method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control.
Grounded expressiveness+2masked-source-output review method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control.
Input readability+1masked-source review method. Blind* absolute scores: 5 for [T] versus 4 for the strongest control.
Output readability+1masked-output review method. Blind* absolute scores: 6 for [T] versus 5 for the strongest control.
Constraint integration+2masked-source-output review method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control.
Reusable abstraction quality+1masked-source review method. Blind* absolute scores: 6 for [T] versus 5 for the strongest control.
Total+11Net contrastive gain/loss.
Metric definitions and scoring legendOpen for exact meanings.

Contrastive score

A -3..+3 judgment comparing [T] against the strongest listed control for each criterion.

Total score

The sum of the seven contrastive criterion scores for one study.

Score color

Green means [T] is better, red means worse, amber means similar; intensity follows magnitude.

What improved and what failed

What improved

  • [T] WeaveMark treatment wins source-only leverage: 18.37 versus 11.36 for [C2] Matched reusable-template control.
  • [T] WeaveMark treatment wins discounted fact units: 289.5 versus 83 for [C2] Matched reusable-template control.
  • The treatment integrates tower defense, deckbuilding, ecosystem simulation, assets, state, balance, UI, and validation into one playable trace.
  • It wins leverage, information yield, and total fact units against the matched template.
  • It is a useful stress test for whether several reusable mechanics can shape one final specification.

What failed or did not improve

  • [T] WeaveMark treatment loses information density: 52 versus 87 for [C2] Matched reusable-template control.
  • [T] WeaveMark treatment loses information yield: 955.4 versus 988.1 for [C2] Matched reusable-template control.
  • [T] WeaveMark treatment is much longer: 5,567 words versus 954 for [C2] Matched reusable-template control.
  • The treatment is much longer and less dense than the matched template.
  • It is a synthetic game concept, so it should not carry the main real-work application claim.
  • No generated browser game has been implemented and tested yet.

Interpretation

A strong structural-mingling stress test, with length/density and synthetic-domain caveats. The qualitative claim should include both sides: WeaveMark improves semantic integration where shown, but the measured failures and caveats are part of the result.