WeaveMark qualitative report

Learning Tutor Final Quality Analysis

A pasteable linear-algebra tutor prompt that teaches through geometric intuition, Socratic questions, misconception diagnosis, adaptive practice, and delayed review.

Outputs inspected

VariantOutputLinesWords
[C1] Compact manual00-control-compact-manual-linear-algebra-tutor.md542
[C2] Matched prose control01-control-matched-prose-linear-algebra-tutor.md35236
[T] WeaveMark treatment02-treatment-refined-expand-linear-algebra-tutor.md2792,066
Metric definitions and scoring legendOpen for exact meanings.

Lines

Saved output line count.

Words

Saved compiled output word count.

Verbatim snippets

Verbatim source/output material quoted from saved artifacts.

[C1] Compact manualOpen output
# Compact Manual Linear Algebra Tutor
[C2] Matched prose controlOpen output
The tutor should teach a motivated returning beginner who remembers algebra but has weak geometric intuition. The tutor should not lecture all at once. It should begin by probing the learner's current mental model, ask one focused question at a time, and adapt based on the answer.
[T] WeaveMark treatmentOpen output
# Linear Algebra Tutor Prompt
[T] WeaveMark treatmentOpen source seam
The final prompt must be one coherent tutor prompt. It should define the tutor role, first interaction, adaptive question sequence, misconception diagnosis, practice ladder, feedback rules, and final mastery check.

Contrastive gain/loss scores

Scores compare [T] WeaveMark treatment against [C2] Matched prose control on the -3..+3 scale.

Primary scores are blind* using hybrid-derived-metrics-and-masked-review: anonymous absolute 1..7 scores were frozen before reveal, then converted to the -3..+3 treatment-control scale. Hybrid blind* scoring uses derived metrics for mechanical criteria and masked source/output review for criteria that require actual reading. The masked review is less blind because domain content, source syntax, or style can leak, but this is necessary to avoid replacing readability and integration judgments with weak length/density proxies.

CriterionBlind* scoreEvidence
Authoring leverage+3derived-evidence method. Blind* absolute scores: 7 for [T] versus 1 for the strongest control.
Information yield+2derived-evidence method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control.
Grounded expressiveness+3masked-source-output review method. Blind* absolute scores: 7 for [T] versus 2 for the strongest control.
Input readability-1masked-source review method. Blind* absolute scores: 5 for [T] versus 6 for the strongest control.
Output readability+2masked-output review method. Blind* absolute scores: 7 for [T] versus 4 for the strongest control.
Constraint integration+3masked-source-output review method. Blind* absolute scores: 7 for [T] versus 2 for the strongest control.
Reusable abstraction quality+2masked-source review method. Blind* absolute scores: 6 for [T] versus 2 for the strongest control.
Total+14Net contrastive gain/loss.
Metric definitions and scoring legendOpen for exact meanings.

Contrastive score

A -3..+3 judgment comparing [T] against the strongest listed control for each criterion.

Total score

The sum of the seven contrastive criterion scores for one study.

Score color

Green means [T] is better, red means worse, amber means similar; intensity follows magnitude.

What improved and what failed

What improved

  • [T] WeaveMark treatment wins source-only leverage: 12.6 versus 1 for [C2] Matched prose control.
  • [T] WeaveMark treatment wins discounted fact units: 141.25 versus 18 for [C2] Matched prose control.
  • [T] WeaveMark treatment wins information yield: 861.3 versus 76.3 for [C2] Matched prose control.
  • The treatment uses fewer local source words than the matched-prose control and produces far more semantic content.
  • Pedagogy, diagnosis, practice, branching, and delayed review become one concrete tutor behavior.
  • This is a strong non-programming demonstration of reusable refinement.

What failed or did not improve

  • [T] WeaveMark treatment loses information density: 68.4 versus 76.3 for [C2] Matched prose control.
  • [T] WeaveMark treatment is much longer: 2,066 words versus 236 for [C2] Matched prose control.
  • The artifact is still a prompt for a tutor, not a measured learner outcome.
  • Because the control is matched prose rather than a reusable-template control, it is not apples-to-apples with headline software-specification studies.
  • The final tutor prompt is much longer, so usability depends on whether the receiving model follows the structure.

Interpretation

A strong supporting non-programming result, especially on leverage and yield versus matched prose. The qualitative claim should include both sides: WeaveMark improves semantic integration where shown, but the measured failures and caveats are part of the result.