Exploratory outcome
9/9 positive
0 neutral, 0 negative deltas in this corpus.
Single-output semantic reuse with structural mingling is the strongest result; matched-template density and matched-prose @expand wins remain visible caveats.
Markers and headline result.
Exploratory outcome
9/9 positive
0 neutral, 0 negative deltas in this corpus.
Net score
+89*
Primary blind* contrastive score.
Study count
9
Single-output study corpus.
Headline studies
4
Realistic/product and key stress-test cases.
Metric rows
27
Every [C1], [C2], and [T] variant.
What a reader should understand after skimming.
+89* aggregate across 9 single-output studies in this corpus: 9 positive deltas, 0 neutral, and 0 negative deltas.
Within this corpus, reusable refinements score best when they create semantic structure that is mingled through one final artifact rather than appended.
Learning Tutor (+14), Evidence-to-Decision Workspace (+11), Orbital Drift (+11), Verdant Relay (+11).
none. Negative deltas occur where matched prose is already concrete or where the treatment gives up density/yield.
The evidence supports reusable specification refinement and structural mingling, not a universal claim that every output is shorter, denser, or behaviorally better.
Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.
Words in the local study source for a variant; this is the local authoring burden.
Words in a variant's explicit input payload, when a template or refinement uses one.
Words in the saved compiled final artifact.
Output words divided by local source words; larger means more final artifact per local word, not quality by itself.
Novelty-weighted semantic fact units extracted from the output by deterministic rules.
Discounted fact units per 1,000 output words; higher means a more information-dense output.
Discounted fact units per 1,000 local source words; higher means more semantic material per local authoring word.
A -3..+3 judgment comparing [T] against the strongest listed control for each criterion.
The sum of the seven contrastive criterion scores for one study.
Green means [T] is better, red means worse, amber means similar; intensity follows magnitude.
An anonymous packet containing automated metrics and extracted fact candidates instead of raw source/output text.
An anonymous 1..7 criterion score assigned before applying the derandomization key; higher is better.
Treatment absolute score minus control absolute score, mapped back to the -3..+3 contrastive scale after reveal.
Remaining direct labels such as [C1], [C2], [T], WeaveMark, control/treatment, or directives after masking.
Deterministic treatment-control metric context after reveal. Bold green numbers are the winner for each metric; red numbers show the losing side. Caveats are summarized separately below.
Leverage
[T] wins
Fact units
[T] wins
Yield
[C2] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[C2] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[T] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[T] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[T] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[T] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[C2] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[T] wins
Leverage
[T] wins
Fact units
[T] wins
Yield
[T] wins
Long findings are kept out of the metric race so the numbers stay scannable.
[T] WeaveMark treatment loses information density: 67.6 versus 91.4 for [C2] Matched reusable-template control.
[T] WeaveMark treatment loses information density: 33 versus 90.8 for [C2] Matched reusable-template control.
[T] WeaveMark treatment loses information density: 59.8 versus 88.4 for [C2] Matched reusable-template control.
[T] WeaveMark treatment loses information density: 68.4 versus 76.3 for [C2] Matched prose control.
[T] WeaveMark treatment loses information density: 66.1 versus 66.3 for [C2] Matched reusable-template control.
[T] WeaveMark treatment loses information density: 65.8 versus 81.8 for [C2] Matched reusable-template control.
[T] WeaveMark treatment loses information density: 52 versus 87 for [C2] Matched reusable-template control.
The matched-prose control remains the fairness baseline because it spells out the same inspiration set without `@expand`.
The source concepts are concrete enough that manual prose can unpack them very effectively.
Scores use the -3..+3 scale. Color intensity is proportional to magnitude; criterion-aware blind* values are primary where available.
| Study and comparison | Authoring leverage | Information yield | Grounded expressiveness | Input readability | Output readability | Constraint integration | Reusable abstraction | Total score |
|---|---|---|---|---|---|---|---|---|
[T]vs[C2] | +2 | -2 | +2 | +1 | +1 | +2 | +1 | +7 |
[T]vs[C2] | +2 | -2 | +2 | +1 | +1 | +2 | +1 | +7 |
[T]vs[C2] | +2 | +2 | +2 | +1 | +1 | +2 | +1 | +11 |
[T]vs[C2] | +3 | +2 | +3 | -1 | +2 | +3 | +2 | +14 |
[T]vs[C2] | +2 | +2 | +1 | +1 | +1 | +1 | +1 | +9 |
[T]vs[C2] | +2 | +2 | +2 | +1 | +1 | +2 | +1 | +11 |
[T]vs[C2] | +2 | +2 | +2 | +1 | +1 | +2 | +1 | +11 |
[T]vs[C2] | +3 | +3 | +1 | -1 | +0 | +1 | +1 | +8 |
[T]vs[C2] | +3 | +3 | +2 | -1 | +1 | +2 | +1 | +11 |
Fact-unit totals are deterministic semantic-information proxies, not behavioral proof.
Headline controls
374
284 local words.
Headline treatments
1,145.75
1,019 local words.
All controls
653.25
8,145 output words.
All treatments
2,087
34,618 output words.
Anonymous scores are frozen before reveal; the asterisk marks the remaining criterion-specific leakage caveat.
Primary blind* delta
+89*
Used as the primary score source where available.
Run
20260710T1416-iterate-final
source-and-output / hybrid-derived-metrics-and-masked-review
Packets
27
0 direct marker leaks after masking.
| Study | Blind* delta |
|---|---|
| Crowd Factory Puzzle | +11 |
| Evidence-to-Decision Workspace | +11 |
| Learning Tutor | +14 |
| Orbital Drift | +11 |
| Release Readiness Workbench | +7 |
| Research Brief | +9 |
| Transit City Swarm | +8 |
| Verdant Relay | +11 |
| Intelligence-to-Execution Kanban | +7 |
*Leakage risk note: Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.
| Criterion | Blindness level | Method | Leakage risk |
|---|---|---|---|
| Authoring leverage | derived-evidence | Ranked from local leverage without exposing variant identity. | Low: uses automated counts and ratios only. |
| Constraint integration | masked-source-output review | Independent review reads masked source/output to judge whether constraints are woven into the artifact. | Moderate: domain content and artifact structure may leak. |
| Grounded expressiveness | masked-source-output review | Independent review reads masked source/output because richness and grounding are semantic judgments. | Moderate: domain content and style may leak even after identity masking. |
| Information yield | derived-evidence | Ranked from discounted fact units per local source word. | Low-to-moderate: fact extraction reads artifacts, but scoring uses derived counts. |
| Input readability | masked-source review | Independent review reads masked source text because readability is not reducible to source length. | Moderate: source syntax/style may reveal the authoring approach. |
| Output readability | masked-output review | Independent review reads masked output text because readability is not reducible to density or brevity. | Moderate: output style/domain content may leak. |
| Reusable abstraction quality | masked-source review | Independent review reads masked source to judge abstraction clarity and reuse. | Moderate-to-high: abstraction syntax can leak authoring style, but reading it is necessary for reliability. |
A local-first release command center that turns release notes, docs, validation runs, screenshots, package artifacts, risks, waivers, and go/no-go decisions into one auditable workspace.
+7 Net contrastive score
A local-first Kanban board for monitoring selected topics, turning signals into cards, deciding actions, delegating work, tracking status, and preserving output lineage.
+7 Net contrastive score
A local-first analyst workspace that turns documents, notes, links, news, claims, contradictions, options, decisions, and follow-up actions into an auditable decision surface.
+11 Net contrastive score
A pasteable linear-algebra tutor prompt that teaches through geometric intuition, Socratic questions, misconception diagnosis, adaptive practice, and delayed review.
+14 Net contrastive score
A concise research-brief instruction for energy-storage strategy that requires source families, context limits, contradictions, alternatives, caveats, and explainable evidence handling.
+9 Net contrastive score
A browser racing game about piloting a small craft through asteroid fields, gravity wells, orbital gates, lap routing, hazards, scoring, restart, and browser validation.
+11 Net contrastive score
A browser game about defending a living railway garden from blight by combining tower-defense route pressure, deckbuilder card choices, ecosystem feedback, original assets, and browser validation.
+11 Net contrastive score
A browser strategy game that combines transit-network drawing, city growth, and ant-colony pathfinding through pheromone-style demand trails and congestion feedback.
+8 Net contrastive score
A browser puzzle game about steering autonomous crowds through factory automation, belts, machines, crates, spatial pushing rules, hazards, and readable level constraints.
+11 Net contrastive score
Best gain: [T] WeaveMark treatment wins source-only leverage: 19.2 versus 15.72 for [C2] Matched reusable-template control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 67.6 versus 91.4 for [C2] Matched reusable-template control.
Conclusion: A strong headline study, with the honest caveat that the template remains denser and more source-efficient on the yield proxy.
Best gain: [T] WeaveMark treatment wins source-only leverage: 19.07 versus 16.4 for [C2] Matched reusable-template control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 33 versus 90.8 for [C2] Matched reusable-template control.
Conclusion: A strong realistic study for semantic propagation, with a measured density/yield loss that should stay visible.
Best gain: [T] WeaveMark treatment wins source-only leverage: 26.63 versus 16.3 for [C2] Matched reusable-template control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 59.8 versus 88.4 for [C2] Matched reusable-template control.
Conclusion: The strongest realistic application result on total semantic content and yield, though not on compactness.
Best gain: [T] WeaveMark treatment wins source-only leverage: 12.6 versus 1 for [C2] Matched prose control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 68.4 versus 76.3 for [C2] Matched prose control.
Conclusion: A strong supporting non-programming result, especially on leverage and yield versus matched prose.
Best gain: [T] WeaveMark treatment wins source-only leverage: 8.34 versus 8.1 for [C2] Matched reusable-template control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 66.1 versus 66.3 for [C2] Matched reusable-template control.
Conclusion: A modest but realistic supporting win whose value is quality-lens integration more than raw metric dominance.
Best gain: [T] WeaveMark treatment wins source-only leverage: 22.81 versus 11.02 for [C2] Matched reusable-template control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 65.8 versus 81.8 for [C2] Matched reusable-template control.
Conclusion: A strong game-specification result, best used as supporting implementation-spec evidence rather than the main claim.
Best gain: [T] WeaveMark treatment wins source-only leverage: 18.37 versus 11.36 for [C2] Matched reusable-template control.
Important failure/caveat: [T] WeaveMark treatment loses information density: 52 versus 87 for [C2] Matched reusable-template control.
Conclusion: A strong structural-mingling stress test, with length/density and synthetic-domain caveats.
Best gain: [T] Expanded WeaveMark treatment wins source-only leverage: 16.36 versus 7.83 for [C2] Matched-prose no-expand control.
Important failure/caveat: The matched-prose control remains the fairness baseline because it spells out the same inspiration set without `@expand`.
Conclusion: A useful `@expand` study where compact named inspirations now produce stronger deterministic proxy metrics than matched prose, while still needing behavioral proof.
Best gain: [T] Expanded WeaveMark treatment wins source-only leverage: 17.46 versus 7.22 for [C2] Matched-prose no-expand control.
Important failure/caveat: The source concepts are concrete enough that manual prose can unpack them very effectively.
Conclusion: A positive `@expand` result: useful for clarity and framing, and currently ahead of matched prose on deterministic proxy metrics.