WeaveMark study corpus

WeaveMark Studies Result

Single-output semantic reuse with structural mingling is the strongest result; matched-template density and matched-prose @expand wins remain visible caveats.

At a glance

Markers and headline result.

[C1]Compact
[C2]Matched
[T]WeaveMark

Exploratory outcome

9/9 positive

0 neutral, 0 negative deltas in this corpus.

Net score

+89*

Primary blind* contrastive score.

Study count

9

Single-output study corpus.

Headline studies

4

Realistic/product and key stress-test cases.

Metric rows

27

Every [C1], [C2], and [T] variant.

Key insights

What a reader should understand after skimming.

Primary score signal

+89* aggregate across 9 single-output studies in this corpus: 9 positive deltas, 0 neutral, and 0 negative deltas.

Pattern in positive deltas

Within this corpus, reusable refinements score best when they create semantic structure that is mingled through one final artifact rather than appended.

Largest positive deltas

Learning Tutor (+14), Evidence-to-Decision Workspace (+11), Orbital Drift (+11), Verdant Relay (+11).

Negative deltas

none. Negative deltas occur where matched prose is already concrete or where the treatment gives up density/yield.

Honest claim

The evidence supports reusable specification refinement and structural mingling, not a universal claim that every output is shorter, denser, or behaviorally better.

Blindness caveat*

Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.

Metric definitions and scoring legendOpen for exact meanings.

Source words

Words in the local study source for a variant; this is the local authoring burden.

Variable words

Words in a variant's explicit input payload, when a template or refinement uses one.

Output words

Words in the saved compiled final artifact.

Leverage

Output words divided by local source words; larger means more final artifact per local word, not quality by itself.

Fact units

Novelty-weighted semantic fact units extracted from the output by deterministic rules.

Density

Discounted fact units per 1,000 output words; higher means a more information-dense output.

Yield

Discounted fact units per 1,000 local source words; higher means more semantic material per local authoring word.

Contrastive score

A -3..+3 judgment comparing [T] against the strongest listed control for each criterion.

Total score

The sum of the seven contrastive criterion scores for one study.

Score color

Green means [T] is better, red means worse, amber means similar; intensity follows magnitude.

Derived-evidence packet

An anonymous packet containing automated metrics and extracted fact candidates instead of raw source/output text.

Blind absolute score

An anonymous 1..7 criterion score assigned before applying the derandomization key; higher is better.

Blind delta

Treatment absolute score minus control absolute score, mapped back to the -3..+3 contrastive scale after reveal.

Direct marker leaks

Remaining direct labels such as [C1], [C2], [T], WeaveMark, control/treatment, or directives after masking.

Post-reveal metric race

Deterministic treatment-control metric context after reveal. Bold green numbers are the winner for each metric; red numbers show the losing side. Caveats are summarized separately below.

[C2][T]

Leverage

[C2]15.72x
[T]19.2x

[T] wins

Fact units

[C2]97.75
[T]322

[T] wins

Yield

[C2]1,437.5
[T]1,298.4

[C2] wins

[C2][T]

Leverage

[C2]1x
[T]12.6x

[T] wins

Fact units

[C2]18
[T]141.25

[T] wins

Yield

[C2]76.3
[T]861.3

[T] wins

[C2][T]

Leverage

[C2]8.1x
[T]8.34x

[T] wins

Fact units

[C2]22
[T]113.5

[T] wins

Yield

[C2]536.6
[T]551

[T] wins

[C2][T]

Leverage

[C2]11.02x
[T]22.81x

[T] wins

Fact units

[C2]55
[T]228.25

[T] wins

Yield

[C2]901.6
[T]1,501.6

[T] wins

[C2][T]

Leverage

[C2]11.36x
[T]18.37x

[T] wins

Fact units

[C2]83
[T]289.5

[T] wins

Yield

[C2]988.1
[T]955.4

[C2] wins

[C2][T]

Leverage

[C2]7.83x
[T]16.36x

[T] wins

Fact units

[C2]95
[T]213.5

[T] wins

Yield

[C2]527.8
[T]1,206.2

[T] wins

[C2][T]

Leverage

[C2]7.22x
[T]17.46x

[T] wins

Fact units

[C2]89.25
[T]244.75

[T] wins

Yield

[C2]490.4
[T]1,281.4

[T] wins

Key caveats to notice

Long findings are kept out of the metric race so the numbers stay scannable.

Release Readiness Workbench · Density loss

[T] WeaveMark treatment loses information density: 67.6 versus 91.4 for [C2] Matched reusable-template control.

Intelligence-to-Execution Kanban · Density loss

[T] WeaveMark treatment loses information density: 33 versus 90.8 for [C2] Matched reusable-template control.

Evidence-to-Decision Workspace · Density loss

[T] WeaveMark treatment loses information density: 59.8 versus 88.4 for [C2] Matched reusable-template control.

Learning Tutor · Density loss

[T] WeaveMark treatment loses information density: 68.4 versus 76.3 for [C2] Matched prose control.

Research Brief · Density loss

[T] WeaveMark treatment loses information density: 66.1 versus 66.3 for [C2] Matched reusable-template control.

Orbital Drift · Density loss

[T] WeaveMark treatment loses information density: 65.8 versus 81.8 for [C2] Matched reusable-template control.

Verdant Relay · Density loss

[T] WeaveMark treatment loses information density: 52 versus 87 for [C2] Matched reusable-template control.

Transit City Swarm · The matched-prose control remains the fairness b

The matched-prose control remains the fairness baseline because it spells out the same inspiration set without `@expand`.

Crowd Factory Puzzle · The source concepts are concrete enough that man

The source concepts are concrete enough that manual prose can unpack them very effectively.

Contrastive gain/loss scores

Scores use the -3..+3 scale. Color intensity is proportional to magnitude; criterion-aware blind* values are primary where available.

Study and comparisonAuthoring
leverage
Information
yield
Grounded
expressiveness
Input
readability
Output
readability
Constraint
integration
Reusable
abstraction
Total
score
+2-2+2+1+1+2+1+7
[T]vs[C2]
+2-2+2+1+1+2+1+7
+2+2+2+1+1+2+1+11
+3+2+3-1+2+3+2+14
+2+2+1+1+1+1+1+9
[T]vs[C2]
+2+2+2+1+1+2+1+11
[T]vs[C2]
+2+2+2+1+1+2+1+11
[T]vs[C2]
+3+3+1-1+0+1+1+8
[T]vs[C2]
+3+3+2-1+1+2+1+11

Aggregate signal

Fact-unit totals are deterministic semantic-information proxies, not behavioral proof.

Headline controls

374

284 local words.

Headline treatments

1,145.75

1,019 local words.

All controls

653.25

8,145 output words.

All treatments

2,087

34,618 output words.

Primary blind* score source

Anonymous scores are frozen before reveal; the asterisk marks the remaining criterion-specific leakage caveat.

Primary blind* delta

+89*

Used as the primary score source where available.

Run

20260710T1416-iterate-final

source-and-output / hybrid-derived-metrics-and-masked-review

Packets

27

0 direct marker leaks after masking.

StudyBlind* delta
Crowd Factory Puzzle+11
Evidence-to-Decision Workspace+11
Learning Tutor+14
Orbital Drift+11
Release Readiness Workbench+7
Research Brief+9
Transit City Swarm+8
Verdant Relay+11
Intelligence-to-Execution Kanban+7

*Leakage risk note: Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.

Criterion-specific blindness policyOpen to inspect why some criteria require masked reading.
CriterionBlindness levelMethodLeakage risk
Authoring leveragederived-evidenceRanked from local leverage without exposing variant identity.Low: uses automated counts and ratios only.
Constraint integrationmasked-source-output reviewIndependent review reads masked source/output to judge whether constraints are woven into the artifact.Moderate: domain content and artifact structure may leak.
Grounded expressivenessmasked-source-output reviewIndependent review reads masked source/output because richness and grounding are semantic judgments.Moderate: domain content and style may leak even after identity masking.
Information yieldderived-evidenceRanked from discounted fact units per local source word.Low-to-moderate: fact extraction reads artifacts, but scoring uses derived counts.
Input readabilitymasked-source reviewIndependent review reads masked source text because readability is not reducible to source length.Moderate: source syntax/style may reveal the authoring approach.
Output readabilitymasked-output reviewIndependent review reads masked output text because readability is not reducible to density or brevity.Moderate: output style/domain content may leak.
Reusable abstraction qualitymasked-source reviewIndependent review reads masked source to judge abstraction clarity and reuse.Moderate-to-high: abstraction syntax can leak authoring style, but reading it is necessary for reliability.

Study coverage

Release Readiness Workbench

A local-first release command center that turns release notes, docs, validation runs, screenshots, package artifacts, risks, waivers, and go/no-go decisions into one auditable workspace.

+7 Net contrastive score

Learning Tutor

A pasteable linear-algebra tutor prompt that teaches through geometric intuition, Socratic questions, misconception diagnosis, adaptive practice, and delayed review.

+14 Net contrastive score

Research Brief

A concise research-brief instruction for energy-storage strategy that requires source families, context limits, contradictions, alternatives, caveats, and explainable evidence handling.

+9 Net contrastive score

Orbital Drift

A browser racing game about piloting a small craft through asteroid fields, gravity wells, orbital gates, lap routing, hazards, scoring, restart, and browser validation.

+11 Net contrastive score

Verdant Relay

A browser game about defending a living railway garden from blight by combining tower-defense route pressure, deckbuilder card choices, ecosystem feedback, original assets, and browser validation.

+11 Net contrastive score

Transit City Swarm

A browser strategy game that combines transit-network drawing, city growth, and ant-colony pathfinding through pheromone-style demand trails and congestion feedback.

+8 Net contrastive score

Crowd Factory Puzzle

A browser puzzle game about steering autonomous crowds through factory automation, belts, machines, crates, spatial pushing rules, hazards, and readable level constraints.

+11 Net contrastive score

Post-reveal qualitative gains and failures

Release Readiness Workbench

Best gain: [T] WeaveMark treatment wins source-only leverage: 19.2 versus 15.72 for [C2] Matched reusable-template control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 67.6 versus 91.4 for [C2] Matched reusable-template control.

Conclusion: A strong headline study, with the honest caveat that the template remains denser and more source-efficient on the yield proxy.

Intelligence-to-Execution Kanban

Best gain: [T] WeaveMark treatment wins source-only leverage: 19.07 versus 16.4 for [C2] Matched reusable-template control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 33 versus 90.8 for [C2] Matched reusable-template control.

Conclusion: A strong realistic study for semantic propagation, with a measured density/yield loss that should stay visible.

Evidence-to-Decision Workspace

Best gain: [T] WeaveMark treatment wins source-only leverage: 26.63 versus 16.3 for [C2] Matched reusable-template control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 59.8 versus 88.4 for [C2] Matched reusable-template control.

Conclusion: The strongest realistic application result on total semantic content and yield, though not on compactness.

Learning Tutor

Best gain: [T] WeaveMark treatment wins source-only leverage: 12.6 versus 1 for [C2] Matched prose control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 68.4 versus 76.3 for [C2] Matched prose control.

Conclusion: A strong supporting non-programming result, especially on leverage and yield versus matched prose.

Research Brief

Best gain: [T] WeaveMark treatment wins source-only leverage: 8.34 versus 8.1 for [C2] Matched reusable-template control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 66.1 versus 66.3 for [C2] Matched reusable-template control.

Conclusion: A modest but realistic supporting win whose value is quality-lens integration more than raw metric dominance.

Orbital Drift

Best gain: [T] WeaveMark treatment wins source-only leverage: 22.81 versus 11.02 for [C2] Matched reusable-template control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 65.8 versus 81.8 for [C2] Matched reusable-template control.

Conclusion: A strong game-specification result, best used as supporting implementation-spec evidence rather than the main claim.

Verdant Relay

Best gain: [T] WeaveMark treatment wins source-only leverage: 18.37 versus 11.36 for [C2] Matched reusable-template control.

Important failure/caveat: [T] WeaveMark treatment loses information density: 52 versus 87 for [C2] Matched reusable-template control.

Conclusion: A strong structural-mingling stress test, with length/density and synthetic-domain caveats.

Transit City Swarm

Best gain: [T] Expanded WeaveMark treatment wins source-only leverage: 16.36 versus 7.83 for [C2] Matched-prose no-expand control.

Important failure/caveat: The matched-prose control remains the fairness baseline because it spells out the same inspiration set without `@expand`.

Conclusion: A useful `@expand` study where compact named inspirations now produce stronger deterministic proxy metrics than matched prose, while still needing behavioral proof.

Crowd Factory Puzzle

Best gain: [T] Expanded WeaveMark treatment wins source-only leverage: 17.46 versus 7.22 for [C2] Matched-prose no-expand control.

Important failure/caveat: The source concepts are concrete enough that manual prose can unpack them very effectively.

Conclusion: A positive `@expand` result: useful for clarity and framing, and currently ahead of matched prose on deterministic proxy metrics.

What not to claim yet

  • Do not claim downstream users or programming agents perform better; that has not been measured.
  • Do not treat contrastive scores or semantic-information proxies as behavioral proof.
  • Do not treat output length as quality unless added text introduces operational obligations.
  • Do not hide negative results: lower density, lower yield, weaker readability, and matched-prose wins are evidence.