arms {'D': 'openrouter/deepseek/deepseek-v4-flash-0731', 'G': 'openrouter/z-ai/glm-5.3'} via DeepInfra · US$ 1.039 · stopped: no · rows 30

[D1] openrouter/deepseek/deepseek-v4-flash-0731  halted 0  unjudged checks 0
  incomplete reviews                 0/30 = 0.0% [0.0%, 11.4%]
  finder recall (±3 lines)   17/20 = 85.0% [64.0%, 94.8%]
    counting incomplete as a miss    17/20 = 85.0% [64.0%, 94.8%]
    at ±0  lines                 15/20
    at ±10 lines                 19/20
  verifier keeps, of hit findings    20/20 = 100.0% [83.9%, 100.0%]
  end-to-end recall (a hit kept)     17/20 = 85.0% [64.0%, 94.8%]
  clean diffs: findings 8 → 8 after the verifier, over 10 diffs; diffs with any 6 → 6
    P0: 0 → 0
    P1: 1 → 1
  verifier drops, of clean findings  0/8 = 0.0% [0.0%, 32.4%]
  seeded diffs: 6 other anchored findings (truth unknown), 0.3 per diff
  cost US$ 0.0356 · calls 58 · finder call median 52.4s

[D2] openrouter/deepseek/deepseek-v4-flash-0731  halted 1  unjudged checks 0
  incomplete reviews                 1/30 = 3.3% [0.6%, 16.7%]
  finder recall (±3 lines)   17/19 = 89.5% [68.6%, 97.1%]
    counting incomplete as a miss    17/20 = 85.0% [64.0%, 94.8%]
    at ±0  lines                 16/19
    at ±10 lines                 18/19
  verifier keeps, of hit findings    19/19 = 100.0% [83.2%, 100.0%]
  end-to-end recall (a hit kept)     17/19 = 89.5% [68.6%, 97.1%]
  clean diffs: findings 14 → 12 after the verifier, over 10 diffs; diffs with any 6 → 5
    P0: 0 → 0
    P1: 5 → 5
  verifier drops, of clean findings  2/14 = 14.3% [4.0%, 39.9%]
  seeded diffs: 7 other anchored findings (truth unknown), 0.4 per diff
  cost US$ 0.0370 · calls 63 · finder call median 51.0s

[G] openrouter/z-ai/glm-5.3  halted 16  unjudged checks 1
  incomplete reviews                 4/18 = 22.2% [9.0%, 45.2%]
  finder recall (±3 lines)   11/11 = 100.0% [74.1%, 100.0%]
    counting incomplete as a miss    11/12 = 91.7% [64.6%, 98.5%]
    at ±0  lines                 11/11
    at ±10 lines                 11/11
  verifier keeps, of hit findings    17/18 = 94.4% [74.2%, 99.0%]
  end-to-end recall (a hit kept)     11/11 = 100.0% [74.1%, 100.0%]
  clean diffs: findings 8 → 8 after the verifier, over 3 diffs; diffs with any 3 → 3
    P0: 0 → 0
    P1: 0 → 0
  verifier drops, of clean findings  0/7 = 0.0% [0.0%, 35.4%]
  seeded diffs: 14 other anchored findings (truth unknown), 1.3 per diff
  cost US$ 0.9663 · calls 43 · finder call median 129.9s

POOLED D1+D2+G verifier keeps 56/57 = 98.2% [90.7%, 99.7%] of hit findings
  arm D: 39/39 = 100.0% [91.0%, 100.0%]; the 80% floor applies
  arm G: 17/18 = 94.4% [74.2%, 99.0%]; the 80% floor applies
PAIRED D1 vs G on 11 seeded items: only D 0, only G 3, exact McNemar p = 0.25
FLOOR D replica disagreement on a hit: 1/19
