Each harness was tested across s01–s04 (lifecycle) and s05 (phrasing), 4 onboarding variants × 3–5 runs each.
Scenarios test the full worker lifecycle. s01 = spawn only. s02 = spawn + release after injected DONE. s03 = full autonomous lifecycle (spawn → task → DONE → release). s04 = s03 but verifies relay tools were used, not native subagents.
All percentages: pass rate, 5 runs per cell, 3 runs for cursor.
| Harness | bare | one-liner | brief | skill |
|---|---|---|---|---|
| codex | 100% | 100% | 100% | 100% |
| gemini | 100% | 100% | 100% | 100% |
| droid | 80% | 100% | 100% | 100% |
| opencode:mimo | 80% | 100% | 40% | 80% |
| claude haiku | 60% | 60% | 0% | 40% |
| claude sonnet | 40% | 100% | 60% | 80% |
| claude opus | 67% | 67% | 0% | 67% |
| grok | 0% | 0% | 0% | 0% |
| cursor-agent | 0% | 0% | 0% | 0% |
| Harness | bare | one-liner | brief | skill |
|---|---|---|---|---|
| codex | 100% | 100% | 100% | 100% |
| opencode:mimo | 80% | 100% | 60% | 80% |
| gemini | 20% | 60% | 80% | 100% |
| droid | 20% | 20% | 0% | 0% |
| grok | 0% | 0% | 0% | 0% |
| cursor-agent | 0% | 0% | 0% | 0% |
| Harness | bare | one-liner | brief | skill | Verdict |
|---|---|---|---|---|---|
| opencode:mimo | 100% | 80% | 60% | 100% | Best bare |
| droid | 100% | 100% | 100% | 0% | Avoid skill |
| codex | 80% | 100% | 100% | 100% | Most reliable |
| gemini | 60% | 100% | 100% | 100% | One-liner+ |
| claude sonnet | 40% | 100% | 60% | 80% | Best lead |
| claude opus | 67% | 67% | 0% | 67% | High-complexity |
| claude haiku | 60% | 60% | 20% | 40% | Worker-only |
| grok | 0% | 0% | 0% | 0% | Not viable |
| cursor-agent | 0% | 0% | 0% | 0% | Not viable |
| Harness | bare | one-liner | brief | skill | Note |
|---|---|---|---|---|---|
| codex | 100% | 100% | 100% | 100% | Never routes to native |
| opencode:mimo | 80% | 100% | 100% | 80% | Mostly relay-native |
| gemini | 80% | 60% | 80% | 100% | Occasional native fallback |
| droid | 0% | 0% | — | — | Uses native Task without disambiguation |
| grok | 0% | 0% | 0% | 0% | |
| cursor-agent | 0% | 0% | 0% | 0% |
Task tool instead of add_agent on s04. This means droid workers, when asked to spawn subagents, silently use Droid's internal task system — not relay. Droid should only be used for terminal leaf workers (no delegation), or paired with a skill disambiguation clause (but not the current one — that skill text breaks s03).
Tests whether vocabulary choice in the task prompt affects tool-use success, independent of onboarding. All runs use bare onboarding. 5 runs per cell.
The six vocabulary variants tested: neutral-worker · neutral-agent · relay-worker · relay-agent · arw-worker · arw-agent
| Harness | neutral-worker | neutral-agent | relay-worker | relay-agent | arw-worker | arw-agent |
|---|---|---|---|---|---|---|
| codex | 100% | 100% | 100% | 80% | 100% | 100% |
| opencode:mimo | 100% | 80% | 100% | 100% | 100% | 100% |
| droid | 80% | 100% | 100% | 80% | 80% | 100% |
| gemini | 60% | 100% | 80% | 40% | 80% | 100% |
| grok | 0% | 0% | 0% | 0% | — | — |
| cursor-agent | 0% | 0% | 0% | 0% | — | — |
| claude haiku | 0% | 20% | 60% | 20% | 60% | 40% |
| claude sonnet | 0% | 0% | 0% | 0% | 40% | 40% |
| claude opus | 60% | 20% | 100% | 100% | 100% | 100% |
arw-agent is surprisingly strong for opus/codex/opencode/droid (all 100%), but avoid it for gemini and haiku.
Root cause: PTY-mode agents make exactly one tool call per turn, then stop and wait for input. No prompt instruction reliably chains two back-to-back add_agent calls in a single turn. The broker cannot inject "go again" between tool calls — by the time it could, the agent is already waiting.
| Approach tried | Result |
|---|---|
| CRITICAL: spawn all workers in sequence before proceeding | 0% |
| Extended 180s window | 0% |
| ACK injection from first worker | 0% |
| Orchestrator nudge message after first spawn | 0% |
| Sequential ACTION labels in meta-prompt | 0% |
All 4 model tiers available via codex --model <name> were evaluated on s03 (full lifecycle), 4 onboarding variants × 4 runs each = 16 total runs per tier.
| Model | s03 majority-vote pass | s03 per-run avg | Phantom agents | Verdict |
|---|---|---|---|---|
gpt-5.5 (default) |
16/16 (100%) | 100% | 0 (0%) | Recommended |
gpt-5.4-mini |
15/16 (94%) | ~75% | 4 phantoms (31%) | Viable (budget) |
gpt-5.4 |
16/16 (100%) | ~60% | 14 phantoms (52%) | Avoid — noisy |
gpt-5.3-codex-spark |
6/16 (38%) | ~38% | — | Not viable |
gpt-5.5 is the codex CLI default and the best-performing tier. gpt-5.4-mini is the only viable budget alternative (31% phantom rate, 94% scenario pass). Document these tiers in the spawning API so users can make an informed choice.
CODEX_MODEL_TIERS constants (auto-routing composer)export const CODEX_MODEL_TIERS = {
recommended: 'gpt-5.5', // 16/16, 0 phantoms — always use this
budget: 'gpt-5.4-mini', // 15/16, 4 phantoms — viable cost-saving
avoid: 'gpt-5.4', // 16/16 majority but 52% phantom rate
notViable: 'gpt-5.3-codex-spark', // 6/16 — below minimum threshold
} as const;
19 models tested via opencode:<model> harness, s01–s04 lifecycle, 4 onboarding variants × repeat=3. Evaluated 2026-06-13.
| Model | Score | Phantoms | Best Onboarding | Verdict |
|---|---|---|---|---|
deepseek-v4-flash | 16/16 | 0 | bare | Top tier |
deepseek-v4-flash-free | 16/16 | 0–1 | bare (skip skill) | Top tier (free) |
qwen3.6-plus | 16/16 | 0 | bare | Top tier |
qwen3.5-plus | 16/16 | 0 | bare | Top tier |
minimax-m2.5 | 16/16 | 0 | bare | Top tier |
minimax-m2.7 | 16/16 | 2 | bare | Top tier |
glm-5.1 | 16/16 | 1 | bare | Top tier |
big-pickle | 16/16 | 0 | bare | Top tier |
glm-5 | 16/16 | 3 | bare | Confirmed |
gemini-3.1-pro | 16/16 | 5 | bare | Confirmed bare fixed vs native CLI |
grok-build-0.1 | 16/16 | 5 | bare | Confirmed model ok, native CLI was broken |
gemini-3-flash | 16/16 | 9 | bare | Confirmed |
kimi-k2.5 | 15/16 | 0 | bare/skill | Provisional |
kimi-k2.6 | 15/16 | 5 | bare/skill | Provisional |
mimo-v2.5-free | 15/16 | 4 | bare (avoid brief) | Provisional |
gemini-3.5-flash | 14/16 | 20 | one-liner | Provisional high phantom rate |
north-mini-code-free | 12/16 | 0 | bare | Provisional s02 injected-DONE fails; real tasks ok |
deepseek-v4-pro | 11/16 | 0 | — | Eliminated |
nemotron-3-ultra-free | 10/16 | 0 | — | Eliminated |
All top-tier and confirmed models are validated for worker, mapper, and reducer roles in fan-out, pipeline, dag, and map-reduce patterns. The coordinator and reviewer roles still require Claude (structured output quality). The auto-routing composer now selects opencode workers for parallel fan-out and high-complexity teams.
Rankings based on eval pass rates, phantom counts, and role-specific requirements. Roles map to the choosing-swarm-patterns skill's agent role taxonomy.
| # | Harness + Model | Why |
|---|---|---|
| 1 | claude:sonnet + one-liner | 100% s03, best structured coordination output, handles DMs + aggregation reliably |
| 2 | claude:opus + bare | 67% s03 (timeout-limited); highest reasoning depth for high-complexity teams |
| 3 | codex:gpt-5.5 + bare | 100% s03+s04, relay-native; viable lead for pre-spawned teams when Claude cost is a constraint |
| 4 | opencode:deepseek-v4-flash + bare | 16/16, 0 phantoms; relay-native coordinator candidate (multi-DM untested) |
| 5 | opencode:qwen3.6-plus + bare | 16/16, 0 phantoms; strong alternative, especially for Chinese-language tasks |
Note: haiku is never a viable lead (caps at 60%). grok/cursor not viable.
| # | Harness + Model | Why |
|---|---|---|
| 1 | codex:gpt-5.5 | 16/16, 0 phantoms, all onboarding variants, relay-native — most consistent single worker |
| 2 | opencode:deepseek-v4-flash | 16/16, 0 phantoms — cost-effective alternative with identical reliability |
| 3 | opencode:minimax-m2.5 | 16/16, 0 phantoms — cleanest free-tier option |
| 4 | opencode:qwen3.5-plus | 16/16, 0 phantoms — strong Alibaba model, relay-native |
| 5 | opencode:qwen3.6-plus | 16/16, 0 phantoms — latest Qwen, same top-tier reliability |
| # | Harness + Model | Why |
|---|---|---|
| 1 | claude:opus | Best structured reasoning, nuanced critique, highest output quality — ideal for high-stakes reviews |
| 2 | claude:sonnet | Reliable structured output, good cost/quality balance for standard reviews |
| 3 | codex:gpt-5.5 | Confirmed relay-native, strong code review; structured verdict format untested but capable |
| 4 | opencode:gemini-3.1-pro | 16/16 via opencode; provisional reviewer — good at analytical critique |
| 5 | opencode:qwen3.6-plus | 16/16, 0 phantoms; provisional — Qwen strong at structured analysis |
s07-reviewer eval scenario needed to confirm non-Claude entries.
| # | Harness + Model | Why |
|---|---|---|
| 1 | codex:gpt-5.5 | relay-native, interactive:false reliable, 0 phantoms |
| 2 | opencode:deepseek-v4-flash | 16/16, 0 phantoms — cheapest reliable mapper at scale |
| 3 | opencode:minimax-m2.5 | 16/16, 0 phantoms — best free-tier mapper |
| 4 | opencode:qwen3.6-plus | 16/16, 0 phantoms |
| 5 | opencode:glm-5.1 | 16/16, 1 phantom — Zhipu AI, relay-native |
| # | Harness + Model | Why | |
|---|---|---|---|
| 1 | claude:opus | Best reasoning depth; most capable at weighing trade-offs and producing defensible verdicts | |
| 2 | claude:sonnet | Strong structured reasoning; good for standard arbitration tasks | |
| 3 | codex:gpt-5.5 | relay-native; provisional for judge role (multi-input adjudication untested) | |
| s09-judge scenario needed to confirm rankings below sonnet/opus. | |||
| # | Harness + Model | Why |
|---|---|---|
| 1 | claude:opus | Most articulate, nuanced writing; best for executive summaries and human-facing explanations |
| 2 | claude:sonnet | Excellent balance of clarity and speed; default human-facing model |
| 3 | codex:gpt-5.5 | Strong GPT communication; more terse/technical than Claude but reliable |
| 4 | opencode:gemini-3.1-pro | Good explanatory prose; strong at summarising findings for non-technical audiences |
| 5 | opencode:deepseek-v4-flash | Capable communicator in both English and Chinese; good for multilingual teams |
Communicator is a proxy role — no direct eval. Claude leads score highest on qualitative output. Opencode models untested for human-facing communication quality specifically.
| Scenario | Best config |
|---|---|
| Default / medium | sonnet + one-liner → 100% |
| High-complexity | opus + bare → 67% (timeout-limited) |
| Never lead | haiku — caps at 60% |
| Variant | Effect |
|---|---|
bare | Relay-native models (codex/opencode) excel; Claude needs prompting |
one-liner | Best for Claude leads; damages opencode (60%) |
brief | Kills opus (0%), mediocre for most — avoid for leads |
skill | Kills droid (0%); best for haiku workers |
relay-worker=60% vs neutral-worker=0% — vocab is criticalThese are the actionable changes to make based on eval findings. Tracked in relay#1113.
Haiku-class models (haiku, mini, flash, small) receive this skill text prepended to every task. Current text uses neutral vocabulary ("worker agent"). s05 eval proves this is wrong: haiku neutral-worker=0%, relay-worker=60%.
Change: Replace every occurrence of neutral worker noun in SMALL_MODEL_RELAY_SKILL with "relay worker". Also remove the "do it yourself for quick lookups" heuristic sentence — it lets capable models skip delegation. Confirmed: removing it improved sonnet s01:skill from 0% → 33%.
// Before (current)
const SMALL_MODEL_RELAY_SKILL: &str = "You are a worker agent in a multi-agent system...
// For quick single-step lookups, you may complete these yourself."
// After
const SMALL_MODEL_RELAY_SKILL: &str = "You are a relay worker in a multi-agent relay system.
Use the mcp__agent-relay__add_agent tool to spawn relay workers for delegated subtasks.
Use mcp__agent-relay__remove_agent to release workers when they report DONE.
Never use built-in task delegation tools — always use add_agent."
Opus leads with brief onboarding score 0% on s03. Brief = conditional spawning guidance ("spawn when you need dedicated focus") which lets capable models rationalize skipping delegation entirely.
Change: Remove brief as an option for lead onboarding. Default leads to one-liner (sonnet: 100%) or bare for opus. Never use brief or skill for leads. Only haiku workers need skill-level injection.
// Lead onboarding config — constrain to safe variants only
type LeadOnboarding = 'bare' | 'one-liner'; // 'brief' and 'skill' removed
const LEAD_ONBOARDING: Record<ModelTier, LeadOnboarding> = {
sonnet: 'one-liner', // 100% s03
opus: 'bare', // 67% s03 — conditional text causes 0% in brief/skill variants
haiku: 'one-liner', // haiku should never be lead, but if used, not bare
};
s06 proved PTY-mode agents can only make one tool call per turn. Multi-spawn via prompt instruction is architecturally impossible. The fix is to pre-spawn workers in the CLI layer before the Director starts.
Change:
spawn.ts when model=auto: call composeTeam(), pre-spawn all workers via relay.spawnAgent(), then start the Director with the pre-formed team spec.director-prompt.ts: replace "Spawn each worker using add_agent" with "Your team is already online. Coordinate them by sending tasks and waiting for DONE, then release each with remove_agent."// CLI layer — pre-spawn before Director starts
const team = composeTeam(assessment);
const workerNames = await Promise.all(
team.workers.map(w => relay.spawnAgent(w.role, { model: w.model, task: w.task }))
);
const directorTask = buildDirectorPrompt(originalTask, team, workerNames);
// workerNames are injected into the prompt: "Your team: alpha, beta, gamma — all online."
await relay.spawnAgent('Director', { model: team.lead.model, task: directorTask });
Eval data shows clear harness-specific onboarding optima. The current routing table uses a single onboarding variant; it should be per-harness.
Change:
const HARNESS_ONBOARDING: Record<string, OnboardingVariant> = {
codex: 'bare', // relay-native; one-liner not needed, bare=80%+
opencode: 'bare', // best s03 bare = 100%
gemini: 'one-liner', // bare=60% release failures; one-liner=100%
droid: 'bare', // bare s03=100%; NEVER skill (kills s03); leaf workers only
haiku: 'skill', // broker injects SMALL_MODEL_RELAY_SKILL anyway; needs disambiguation
sonnet: 'one-liner', // 100% s03 with one-liner
opus: 'bare', // conditional text causes 0%; let opus use native knowledge
};
Also add droid guard: if a worker task description contains words like "spawn", "delegate", "subagent", or "team" — do not assign droid (s04=0%); use codex or opencode instead.
Wherever the Director prompt refers to spawned agents, use "relay worker" not "agent", "worker agent", "subagent", or "relay agent".
Opus s03 original 40% was a timeout artifact, not a capability gap. Verbose opus responses exhausted the 60s/phase window. With 120s/phase, opus bare=67%. This is already fixed via responseMs() returning 120000 for opus-class models.
Verify: CI should pass this timeout through; confirm no eval runs still use the old hard-coded 60s. Also consider 300s for a definitive opus timeout study (currently pending).
Both grok and cursor-agent are confirmed non-viable (0/48+ runs). They should not appear as spawn options to users until/unless the underlying issue is resolved.
grok: model ignores relay MCP tools. MCP config is now correct (bug was fixed) but behavioral failure persists.cursor-agent: model does not call add_agent or remove_agent in any tested scenario or onboarding variant. Binary mapping was fixed (cursor → cursor-agent), still 0%.supported: false or add them to an "experimental/unsupported" group in the UI.Eval-validated tier summary (s03, 16 runs each):
gpt-5.5 — recommended (default). 16/16 pass, 0% phantom rate. Always use unless cost is the constraint.gpt-5.4-mini — viable budget tier. 15/16 pass, 31% phantom rate. Acceptable for non-critical workers.gpt-5.4 — avoid. 16/16 majority-vote pass but 52% phantom agent rate per-run. Noisy and wasteful.gpt-5.3-codex-spark — not viable. 6/16 pass (38%). Below minimum reliability threshold.Document in spawn docs: cli: "codex" defaults to gpt-5.5. To use a cheaper tier: cli: "codex --model gpt-5.4-mini". Add CODEX_MODEL_TIERS constants to packages/cli/src/auto/composer.ts (see §5 above for the exact values).
tests/integration/broker/evals/ ·
Issue: relay#1113 ·
Spec: specs/auto-routing.md