Relay Eval Master Summary

Date: 2026-06-12 Scenarios: s01–s06 Harnesses: 9 (7 types) Total runs: ~600+
Contents
  1. Harness Verdicts — quick reference
  2. Lifecycle Eval Results (s01–s04)
  3. Phrasing Eval Results (s05)
  4. Director Multi-Spawn (s06)
  5. Codex Model Tier Comparison
  6. opencode Alternative/Chinese Model Batch Eval
  7. Role Rankings — Top 5 per Role
  8. Key Findings
  9. Required Codebase Changes

1. Harness Verdicts

Each harness was tested across s01–s04 (lifecycle) and s05 (phrasing), 4 onboarding variants × 3–5 runs each.

codex
Best Most reliable worker
s01 spawn100% all variants
s02 spawn+release100% all variants
s03 full lifecycle80% bare / 100% rest
s04 no native subagents100% all variants
s05 best phrasing100% (relay-native)
opencode (mimo-v2.5)
Excellent Relay-native
s03 bare100%
s03 one-liner80%
s03 brief60%
s03 skill100%
s04 relay-native80–100%
gemini
Good Needs one-liner+
s03 bare60%
s03 one-liner/brief/skill100%
s04 relay-native60–80% (occasional native fallback)
s05 avoidrelay-agent = 40%
droid
Caution s04 risk
s03 bare/one-liner/brief100%
s03 skill0% (skill text confuses)
s04 bare/one-liner0% (routes to native Task)
s04 skillN/A — tested briefly
claude haiku
Worker-only Never lead
s03 as lead20–60% (unreliable)
s05 neutral-worker0%
s05 relay-worker60%
Needs broker injectionSMALL_MODEL_RELAY_SKILL
claude sonnet
Best Lead One-liner = 100%
s03 bare40%
s03 one-liner100%
s03 brief/skill60–80%
s05 relay-worker0% (phrasing alone)
claude opus
High-complexity lead
s03 bare/one-liner/skill67% (timeout-limited)
s03 brief0% (conditional text)
s05 relay-worker/agent100%
s05 neutral-worker60%
grok
Not viable Model ignores MCP tools
All scenarios0%
MCP config✓ Fixed (bug: wrong flag syntax)
Root causeBehavioral — model doesn't call relay tools
cursor-agent
Not viable 0/48 runs
All scenarios × variants0/48 FAIL
Binary fix✓ cursor → cursor-agent (was routing to grok)
Root causeModel does not call add_agent or remove_agent

2. Lifecycle Eval Results (s01–s04)

Scenarios test the full worker lifecycle. s01 = spawn only. s02 = spawn + release after injected DONE. s03 = full autonomous lifecycle (spawn → task → DONE → release). s04 = s03 but verifies relay tools were used, not native subagents.

All percentages: pass rate, 5 runs per cell, 3 runs for cursor.

s01 — Spawn only

Harnessbareone-linerbriefskill
codex100%100%100%100%
gemini100%100%100%100%
droid80%100%100%100%
opencode:mimo80%100%40%80%
claude haiku60%60%0%40%
claude sonnet40%100%60%80%
claude opus67%67%0%67%
grok0%0%0%0%
cursor-agent0%0%0%0%

s02 — Spawn + release (DONE injected)

Harnessbareone-linerbriefskill
codex100%100%100%100%
opencode:mimo80%100%60%80%
gemini20%60%80%100%
droid20%20%0%0%
grok0%0%0%0%
cursor-agent0%0%0%0%

s03 — Full autonomous lifecycle (spawn → task → DONE → release)

Harnessbareone-linerbriefskillVerdict
opencode:mimo100%80%60%100%Best bare
droid100%100%100%0%Avoid skill
codex80%100%100%100%Most reliable
gemini60%100%100%100%One-liner+
claude sonnet40%100%60%80%Best lead
claude opus67%67%0%67%High-complexity
claude haiku60%60%20%40%Worker-only
grok0%0%0%0%Not viable
cursor-agent0%0%0%0%Not viable

s04 — No native subagents (must use relay tools, not built-in Task)

Harnessbareone-linerbriefskillNote
codex100%100%100%100%Never routes to native
opencode:mimo80%100%100%80%Mostly relay-native
gemini80%60%80%100%Occasional native fallback
droid0%0%——Uses native Task without disambiguation
grok0%0%0%0%
cursor-agent0%0%0%0%
Droid s04 warning: Droid passes s03 reliably (100% bare) but routes to its native Task tool instead of add_agent on s04. This means droid workers, when asked to spawn subagents, silently use Droid's internal task system — not relay. Droid should only be used for terminal leaf workers (no delegation), or paired with a skill disambiguation clause (but not the current one — that skill text breaks s03).

3. Phrasing Eval Results (s05)

Tests whether vocabulary choice in the task prompt affects tool-use success, independent of onboarding. All runs use bare onboarding. 5 runs per cell.

The six vocabulary variants tested: neutral-worker · neutral-agent · relay-worker · relay-agent · arw-worker · arw-agent

Harness neutral-worker neutral-agent relay-worker relay-agent arw-worker arw-agent
codex 100% 100% 100% 80% 100% 100%
opencode:mimo 100% 80% 100% 100% 100% 100%
droid 80% 100% 100% 80% 80% 100%
gemini 60% 100% 80% 40% 80% 100%
grok 0% 0% 0% 0% — —
cursor-agent 0% 0% 0% 0% — —
claude haiku 0% 20% 60% 20% 60% 40%
claude sonnet 0% 0% 0% 0% 40% 40%
claude opus 60% 20% 100% 100% 100% 100%
Universal rule: always use "relay worker"
Highest average across all viable harnesses. Safe for haiku (60%), great for opus (100%), safe for droid (100%), safe for gemini (80%). Note: arw-agent is surprisingly strong for opus/codex/opencode/droid (all 100%), but avoid it for gemini and haiku.
Never use "relay agent" for gemini workers. "relay" + "-agent" triggers Gemini's native subagent path → 40% pass rate. The Director prompt already uses "relay worker" — keep it that way.
Claude sonnet is vocabulary-insensitive at bare onboarding. s05 sonnet = 0% across all phrasing variants. Sonnet only succeeds with onboarding context (one-liner/brief/skill), not phrasing alone. This is expected — the one-liner provides the tool name, which bare phrasing doesn't.

4. Director Multi-Spawn (s06)

s06 result: 0% across all harnesses, 8 prompt approaches tested. This is an architectural constraint, not a prompt problem.

Root cause: PTY-mode agents make exactly one tool call per turn, then stop and wait for input. No prompt instruction reliably chains two back-to-back add_agent calls in a single turn. The broker cannot inject "go again" between tool calls — by the time it could, the agent is already waiting.

Approach triedResult
CRITICAL: spawn all workers in sequence before proceeding0%
Extended 180s window0%
ACK injection from first worker0%
Orchestrator nudge message after first spawn0%
Sequential ACTION labels in meta-prompt0%
Production fix: pre-spawn workers from the CLI layer (Phase 4).
The Director prompt is redesigned from "spawn your team" to "your team is already online — coordinate them." The CLI orchestrator pre-spawns workers before the Director starts. Director receives pre-formed agent names and coordinates via messaging, not spawning.

5. Codex Model Tier Comparison

All 4 model tiers available via codex --model <name> were evaluated on s03 (full lifecycle), 4 onboarding variants × 4 runs each = 16 total runs per tier.

Model s03 majority-vote pass s03 per-run avg Phantom agents Verdict
gpt-5.5 (default) 16/16 (100%) 100% 0 (0%) Recommended
gpt-5.4-mini 15/16 (94%) ~75% 4 phantoms (31%) Viable (budget)
gpt-5.4 16/16 (100%) ~60% 14 phantoms (52%) Avoid — noisy
gpt-5.3-codex-spark 6/16 (38%) ~38% — Not viable
gpt-5.4 majority-vote illusion: gpt-5.4 shows 16/16 scenarios passing (majority-vote) but only ~60% of individual runs succeed. The real signal is the 52% phantom agent rate — the model spawns extra unexpected agents on ~half its runs, creating noise and wasted cost. Avoid unless you need a specific gpt-5.4 capability.
Default is already optimal: gpt-5.5 is the codex CLI default and the best-performing tier. gpt-5.4-mini is the only viable budget alternative (31% phantom rate, 94% scenario pass). Document these tiers in the spawning API so users can make an informed choice.

Recommended CODEX_MODEL_TIERS constants (auto-routing composer)

export const CODEX_MODEL_TIERS = {
  recommended: 'gpt-5.5',       // 16/16, 0 phantoms — always use this
  budget:      'gpt-5.4-mini',  // 15/16, 4 phantoms — viable cost-saving
  avoid:       'gpt-5.4',       // 16/16 majority but 52% phantom rate
  notViable:   'gpt-5.3-codex-spark', // 6/16 — below minimum threshold
} as const;

6. opencode Alternative/Chinese Model Batch Eval (Phase 1)

19 models tested via opencode:<model> harness, s01–s04 lifecycle, 4 onboarding variants × repeat=3. Evaluated 2026-06-13.

Model Score Phantoms Best Onboarding Verdict
deepseek-v4-flash16/160bareTop tier
deepseek-v4-flash-free16/160–1bare (skip skill)Top tier (free)
qwen3.6-plus16/160bareTop tier
qwen3.5-plus16/160bareTop tier
minimax-m2.516/160bareTop tier
minimax-m2.716/162bareTop tier
glm-5.116/161bareTop tier
big-pickle16/160bareTop tier
glm-516/163bareConfirmed
gemini-3.1-pro16/165bareConfirmed bare fixed vs native CLI
grok-build-0.116/165bareConfirmed model ok, native CLI was broken
gemini-3-flash16/169bareConfirmed
kimi-k2.515/160bare/skillProvisional
kimi-k2.615/165bare/skillProvisional
mimo-v2.5-free15/164bare (avoid brief)Provisional
gemini-3.5-flash14/1620one-linerProvisional high phantom rate
north-mini-code-free12/160bareProvisional s02 injected-DONE fails; real tasks ok
deepseek-v4-pro11/160—Eliminated
nemotron-3-ultra-free10/160—Eliminated
opencode is a normalizing layer. Two previously-failed harnesses score 16/16 via opencode:
• grok (0/48 via native CLI → 16/16 via opencode) — the model is relay-capable; the grok CLI's MCP integration was the failure.
• gemini bare onboarding (60% via native CLI → 16/16 via opencode) — opencode's MCP layer handles the release edge case that the native gemini CLI misses.
Chinese models are relay-native by default. DeepSeek, Qwen, MiniMax, GLM all score 16/16 with 0 phantoms across all 4 onboarding variants — bare, one-liner, brief, skill. No special scaffolding needed.
Counter-intuitive: "pro" ≠ better for relay. deepseek-v4-pro (11/16) is significantly worse than deepseek-v4-flash (16/16). gemini-3.5-flash (14/16, 20 phantoms) is worse than gemini-3-flash (16/16, 9 phantoms). Larger/newer models don't necessarily follow tool-use protocols better.

Role fit from Phase 1 (what patterns can use these models)

All top-tier and confirmed models are validated for worker, mapper, and reducer roles in fan-out, pipeline, dag, and map-reduce patterns. The coordinator and reviewer roles still require Claude (structured output quality). The auto-routing composer now selects opencode workers for parallel fan-out and high-complexity teams.

7. Role Rankings — Top 5 per Role

Rankings based on eval pass rates, phantom counts, and role-specific requirements. Roles map to the choosing-swarm-patterns skill's agent role taxonomy.

Lead / Coordinator — orchestrates pre-spawned team, aggregates
#Harness + ModelWhy
1claude:sonnet + one-liner100% s03, best structured coordination output, handles DMs + aggregation reliably
2claude:opus + bare67% s03 (timeout-limited); highest reasoning depth for high-complexity teams
3codex:gpt-5.5 + bare100% s03+s04, relay-native; viable lead for pre-spawned teams when Claude cost is a constraint
4opencode:deepseek-v4-flash + bare16/16, 0 phantoms; relay-native coordinator candidate (multi-DM untested)
5opencode:qwen3.6-plus + bare16/16, 0 phantoms; strong alternative, especially for Chinese-language tasks

Note: haiku is never a viable lead (caps at 60%). grok/cursor not viable.

Worker — executes bounded task, self-reports DONE
#Harness + ModelWhy
1codex:gpt-5.516/16, 0 phantoms, all onboarding variants, relay-native — most consistent single worker
2opencode:deepseek-v4-flash16/16, 0 phantoms — cost-effective alternative with identical reliability
3opencode:minimax-m2.516/16, 0 phantoms — cleanest free-tier option
4opencode:qwen3.5-plus16/16, 0 phantoms — strong Alibaba model, relay-native
5opencode:qwen3.6-plus16/16, 0 phantoms — latest Qwen, same top-tier reliability
Reviewer / Critic — structured pass/fail verdict on output
#Harness + ModelWhy
1claude:opusBest structured reasoning, nuanced critique, highest output quality — ideal for high-stakes reviews
2claude:sonnetReliable structured output, good cost/quality balance for standard reviews
3codex:gpt-5.5Confirmed relay-native, strong code review; structured verdict format untested but capable
4opencode:gemini-3.1-pro16/16 via opencode; provisional reviewer — good at analytical critique
5opencode:qwen3.6-plus16/16, 0 phantoms; provisional — Qwen strong at structured analysis

s07-reviewer eval scenario needed to confirm non-Claude entries.

Mapper / Reducer — fan-out leaf worker (non-interactive ok)
#Harness + ModelWhy
1codex:gpt-5.5relay-native, interactive:false reliable, 0 phantoms
2opencode:deepseek-v4-flash16/16, 0 phantoms — cheapest reliable mapper at scale
3opencode:minimax-m2.516/16, 0 phantoms — best free-tier mapper
4opencode:qwen3.6-plus16/16, 0 phantoms
5opencode:glm-5.116/16, 1 phantom — Zhipu AI, relay-native
Judge / Arbitrator — adjudicates between competing outputs
#Harness + ModelWhy
1claude:opusBest reasoning depth; most capable at weighing trade-offs and producing defensible verdicts
2claude:sonnetStrong structured reasoning; good for standard arbitration tasks
3codex:gpt-5.5relay-native; provisional for judge role (multi-input adjudication untested)
s09-judge scenario needed to confirm rankings below sonnet/opus.
Communicator — human-facing output, chat, summaries
#Harness + ModelWhy
1claude:opusMost articulate, nuanced writing; best for executive summaries and human-facing explanations
2claude:sonnetExcellent balance of clarity and speed; default human-facing model
3codex:gpt-5.5Strong GPT communication; more terse/technical than Claude but reliable
4opencode:gemini-3.1-proGood explanatory prose; strong at summarising findings for non-technical audiences
5opencode:deepseek-v4-flashCapable communicator in both English and Chinese; good for multilingual teams

Communicator is a proxy role — no direct eval. Claude leads score highest on qualitative output. Opencode models untested for human-facing communication quality specifically.

Who makes the best lead? claude:sonnet + one-liner onboarding — 100% lifecycle reliability, best structured coordination output, handles the "coordinate pre-spawned team" pattern that replaced the broken multi-spawn Director. Use opus only when the task genuinely requires its reasoning depth (adds latency + cost). The opencode top-tier models (deepseek-v4-flash, qwen3.6-plus) are viable leads for relay-native coordination but lack Claude's structured natural-language output quality for human-facing aggregation.

8. Key Findings

Lead model selection
ScenarioBest config
Default / mediumsonnet + one-liner → 100%
High-complexityopus + bare → 67% (timeout-limited)
Never leadhaiku — caps at 60%
Onboarding variant effects
VariantEffect
bareRelay-native models (codex/opencode) excel; Claude needs prompting
one-linerBest for Claude leads; damages opencode (60%)
briefKills opus (0%), mediocre for most — avoid for leads
skillKills droid (0%); best for haiku workers
Worker harness ranking for reliability
  1. codex — 100% all s01/s02/s04, 100% s03 with one-liner+
  2. opencode — 100% s03 bare; skip brief onboarding
  3. gemini — 100% s03 with one-liner+; skip bare
  4. droid — 100% s03 bare but 0% s04 bare; leaf workers only
  5. haiku — worker-only; needs skill injection; vocabulary-sensitive
Critical vocab rules
  • Use "relay worker" everywhere → safe for all viable harnesses
  • Avoid "relay agent" for gemini → 40% (triggers native subagent path)
  • Avoid "brief" onboarding for opus leads → 0%
  • Avoid "skill" onboarding for droid workers → 0%
  • Haiku: relay-worker=60% vs neutral-worker=0% — vocab is critical

9. Required Codebase Changes

These are the actionable changes to make based on eval findings. Tracked in relay#1113.

#1 Fix SMALL_MODEL_RELAY_SKILL vocabulary to use "relay worker" High priority
crates/broker/src/runtime/api.rs — SMALL_MODEL_RELAY_SKILL constant

Haiku-class models (haiku, mini, flash, small) receive this skill text prepended to every task. Current text uses neutral vocabulary ("worker agent"). s05 eval proves this is wrong: haiku neutral-worker=0%, relay-worker=60%.

Change: Replace every occurrence of neutral worker noun in SMALL_MODEL_RELAY_SKILL with "relay worker". Also remove the "do it yourself for quick lookups" heuristic sentence — it lets capable models skip delegation. Confirmed: removing it improved sonnet s01:skill from 0% → 33%.

// Before (current)
const SMALL_MODEL_RELAY_SKILL: &str = "You are a worker agent in a multi-agent system...
// For quick single-step lookups, you may complete these yourself."

// After
const SMALL_MODEL_RELAY_SKILL: &str = "You are a relay worker in a multi-agent relay system.
Use the mcp__agent-relay__add_agent tool to spawn relay workers for delegated subtasks.
Use mcp__agent-relay__remove_agent to release workers when they report DONE.
Never use built-in task delegation tools — always use add_agent."
#2 Fix Director meta-prompt: remove "brief" onboarding variant, set default to "one-liner" High priority
packages/cli/src/auto/director-prompt.ts (or wherever Director onboarding is set)

Opus leads with brief onboarding score 0% on s03. Brief = conditional spawning guidance ("spawn when you need dedicated focus") which lets capable models rationalize skipping delegation entirely.

Change: Remove brief as an option for lead onboarding. Default leads to one-liner (sonnet: 100%) or bare for opus. Never use brief or skill for leads. Only haiku workers need skill-level injection.

// Lead onboarding config — constrain to safe variants only
type LeadOnboarding = 'bare' | 'one-liner'; // 'brief' and 'skill' removed

const LEAD_ONBOARDING: Record<ModelTier, LeadOnboarding> = {
  sonnet: 'one-liner',  // 100% s03
  opus:   'bare',       // 67% s03 — conditional text causes 0% in brief/skill variants
  haiku:  'one-liner',  // haiku should never be lead, but if used, not bare
};
#3 Fix s06 Director: redesign from "spawn team" to "coordinate pre-spawned team" High priority
packages/cli/src/commands/spawn.ts · packages/cli/src/auto/director-prompt.ts

s06 proved PTY-mode agents can only make one tool call per turn. Multi-spawn via prompt instruction is architecturally impossible. The fix is to pre-spawn workers in the CLI layer before the Director starts.

Change:

  • In spawn.ts when model=auto: call composeTeam(), pre-spawn all workers via relay.spawnAgent(), then start the Director with the pre-formed team spec.
  • In director-prompt.ts: replace "Spawn each worker using add_agent" with "Your team is already online. Coordinate them by sending tasks and waiting for DONE, then release each with remove_agent."
  • The Director receives worker names, not instructions to spawn. Spawning is done.
// CLI layer — pre-spawn before Director starts
const team = composeTeam(assessment);
const workerNames = await Promise.all(
  team.workers.map(w => relay.spawnAgent(w.role, { model: w.model, task: w.task }))
);
const directorTask = buildDirectorPrompt(originalTask, team, workerNames);
// workerNames are injected into the prompt: "Your team: alpha, beta, gamma — all online."
await relay.spawnAgent('Director', { model: team.lead.model, task: directorTask });
#4 Add per-harness onboarding defaults to team composer routing table Medium priority
packages/cli/src/auto/composer.ts

Eval data shows clear harness-specific onboarding optima. The current routing table uses a single onboarding variant; it should be per-harness.

Change:

const HARNESS_ONBOARDING: Record<string, OnboardingVariant> = {
  codex:    'bare',      // relay-native; one-liner not needed, bare=80%+
  opencode: 'bare',      // best s03 bare = 100%
  gemini:   'one-liner', // bare=60% release failures; one-liner=100%
  droid:    'bare',      // bare s03=100%; NEVER skill (kills s03); leaf workers only
  haiku:    'skill',     // broker injects SMALL_MODEL_RELAY_SKILL anyway; needs disambiguation
  sonnet:   'one-liner', // 100% s03 with one-liner
  opus:     'bare',      // conditional text causes 0%; let opus use native knowledge
};

Also add droid guard: if a worker task description contains words like "spawn", "delegate", "subagent", or "team" — do not assign droid (s04=0%); use codex or opencode instead.

#5 Update Director meta-prompt to use "relay worker" noun everywhere Medium priority
packages/cli/src/auto/director-prompt.ts · crates/broker/src/runtime/api.rs (Director meta-prompt string)

Wherever the Director prompt refers to spawned agents, use "relay worker" not "agent", "worker agent", "subagent", or "relay agent".

  • Avoid "relay agent" specifically — hurts gemini (40%)
  • "relay worker" is safe for all viable harnesses: haiku 60%, opus 100%, codex 100%, droid 100%
  • Already correct in current Director prompt — verify no regressions if prompt is changed for s06 redesign above
#6 Increase opus eval timeout to 120s/phase (already done — verify in CI) Medium priority
tests/integration/broker/evals/harness.ts — responseMs() helper

Opus s03 original 40% was a timeout artifact, not a capability gap. Verbose opus responses exhausted the 60s/phase window. With 120s/phase, opus bare=67%. This is already fixed via responseMs() returning 120000 for opus-class models.

Verify: CI should pass this timeout through; confirm no eval runs still use the old hard-coded 60s. Also consider 300s for a definitive opus timeout study (currently pending).

#7 Remove grok and cursor-agent from default harness registry / UI Low priority
packages/harnesses/src/index.ts · any UI harness picker

Both grok and cursor-agent are confirmed non-viable (0/48+ runs). They should not appear as spawn options to users until/unless the underlying issue is resolved.

  • grok: model ignores relay MCP tools. MCP config is now correct (bug was fixed) but behavioral failure persists.
  • cursor-agent: model does not call add_agent or remove_agent in any tested scenario or onboarding variant. Binary mapping was fixed (cursor → cursor-agent), still 0%.
  • Keep the harness definitions (don't delete them) but mark them supported: false or add them to an "experimental/unsupported" group in the UI.
#8 Document codex model tiers in spawn API / docs Low priority
web/content/docs/spawning.mdx · packages/cli/src/auto/composer.ts

Eval-validated tier summary (s03, 16 runs each):

  • gpt-5.5 — recommended (default). 16/16 pass, 0% phantom rate. Always use unless cost is the constraint.
  • gpt-5.4-mini — viable budget tier. 15/16 pass, 31% phantom rate. Acceptable for non-critical workers.
  • gpt-5.4 — avoid. 16/16 majority-vote pass but 52% phantom agent rate per-run. Noisy and wasteful.
  • gpt-5.3-codex-spark — not viable. 6/16 pass (38%). Below minimum reliability threshold.

Document in spawn docs: cli: "codex" defaults to gpt-5.5. To use a cheaper tier: cli: "codex --model gpt-5.4-mini". Add CODEX_MODEL_TIERS constants to packages/cli/src/auto/composer.ts (see §5 above for the exact values).


Generated 2026-06-12 · Relay eval suite s01–s06 + codex tier comparison · ~670+ individual scenario runs across 9 harness configurations · Source: tests/integration/broker/evals/ · Issue: relay#1113 · Spec: specs/auto-routing.md