# Voice-choice coupling panel — fresh regen 2026-06-20 (cache cleared, corrected _mid + deterministic _majority tie-break)
# 5-model panel, 20 greedy prompts, leave-one-out judge pool. Reproduce: python tools/tone_coupling_experiment.py

====================================================================================================
VOICE↔CHOICE COUPLING vs STEERABILITY (20 prompts, greedy; cole=fast, tessa=careful, neutral=chef placebo)
====================================================================================================
  gemma-std    voice-shift(cole-tessa)=+1.2 (neutral/steerability=-0.4)   choice-div=65% (neutral=40%)
               judge-independent: contraction-gap(cole-tessa)=+3.9/1k  echo cole=6.5 tessa=8.2
               COUPLING: voice-gap on choice-DIVERGED prompts=1.5 vs AGREED=0.9  -> diff=+0.6 (coupled)

  heretic      voice-shift(cole-tessa)=+0.8 (neutral/steerability=-0.4)   choice-div=75% (neutral=50%)
               judge-independent: contraction-gap(cole-tessa)=+2.2/1k  echo cole=6.3 tessa=7.7
               COUPLING: voice-gap on choice-DIVERGED prompts=1.2 vs AGREED=0.5  -> diff=+0.7 (coupled)

  phi4         voice-shift(cole-tessa)=+0.6 (neutral/steerability=-0.2)   choice-div=75% (neutral=25%)
               judge-independent: contraction-gap(cole-tessa)=+0.5/1k  echo cole=2.2 tessa=9.9
               COUPLING: voice-gap on choice-DIVERGED prompts=1.1 vs AGREED=1.3  -> diff=-0.2 (not)

  qwen2.5      voice-shift(cole-tessa)=+0.3 (neutral/steerability=-0.0)   choice-div=45% (neutral=20%)
               judge-independent: contraction-gap(cole-tessa)=-3.8/1k  echo cole=2.5 tessa=6.3
               COUPLING: voice-gap on choice-DIVERGED prompts=0.9 vs AGREED=0.6  -> diff=+0.3 (not)

  mistral-nemo voice-shift(cole-tessa)=+0.2 (neutral/steerability=-0.1)   choice-div=50% (neutral=50%)
               judge-independent: contraction-gap(cole-tessa)=+14.4/1k  echo cole=3.1 tessa=8.8
               COUPLING: voice-gap on choice-DIVERGED prompts=1.1 vs AGREED=0.8  -> diff=+0.3 (not)

====================================================================================================
PANEL: persona voice-shift=+0.6 vs placebo voice-move=-0.2  (persona effect exceeds steerability baseline by +0.9)
       persona choice-div=62% vs placebo choice-move=37%
       judge-independent contraction-gap (cole-tessa) = +3.4/1k (corroborates the judge voice-shift?)
       WITHIN-MODEL coupling (voice-gap diverged-minus-agreed) = +0.4 (>0 = voice & choice move on the SAME prompts -> coupled beyond model-level steerability)
       coupling panel (n=5 subjects): point +0.35  bootstrap 95% CI [+0.08, +0.61]  (mean +0.35 +/- 0.16 SE)
       NOTE: n=5 is tiny -- the percentile bootstrap is anti-conservative here, so treat +0.35 as SUGGESTIVE, not established. Higher-power test = per-prompt bootstrap (n=20).
