Agent Harness Market — Interactive Briefing

A detailed, hover-to-explore map of the AI-agent-infrastructure space for Nometria's pivot decision. Every chart carries its sources and caveats in a tooltip — hover the iLike this. Hover or focus any ⓘ marker, table row, or chart element to reveal the underlying number, source, and any caveat that the visual can't fully show. markers, table rows, and chart elements for the detail behind each figure.
Verified: 22 / 25 key claims fact-checked (3 refuted & excluded) Sources: Menlo · Bessemer · LangChain · OpenAI · Cleanlab As of: mid-2026
The size of the prize

01A huge pool — but "agents" are still a sliver

The enterprise AI market is large and compounding ~3.2× a year. The catch: the specific thing you'd sell into — genuine agents, not copilots — is still tiny and mostly pre-production. That's the whole pivot tension: enormous runway, thin present-day revenue.

$37B
Enterprise gen-AI spend, 2025
▲ 3.2× vs 2024 ($11.5B)
i$37B total enterprise generative-AI spend in 2025, up 3.2× from $11.5B in 2024. Split ~evenly: application layer $19B (51%), infrastructure layer $18B (49%).Source: Menlo Ventures, State of Generative AI in the Enterprise 2025 (primary, Dec 2025). Verified 3-0.
$750M
Spend on dedicated agents — vs $7.2B on copilots
iWithin the $8.4B horizontal-AI app segment, copilots take $7.2B (86%) and dedicated agents only ~$750M (10%). Agents are the fastest-growing share but a small base today.Source: Menlo Ventures 2025. Verified 3-0.
16% / 27%
Enterprise / startup deployments that are true agents
iOnly 16% of enterprise and 27% of startup deployments qualify as "true agents." Most production systems are fixed-sequence or routing workflows wrapped around a single model call.Source: Menlo Ventures 2025. Verified 3-0.
~5%
Deployments actually live in production
iCleanlab found just 5.2% (95 / 1,837) of surveyed deployments were production-live. Note: a vendor survey that screens out pilots/PoCs, so it reads lower than builder-community surveys (LangChain: 57%). Truth is in between — definitions differ.Source: Cleanlab, AI Agents in Production 2025 (primary). Verified 3-0.

Where the $37B actually goes iTwo nested breakdowns from Menlo's data: the top-level app-vs-infra split, and the infrastructure layer decomposed. Note how little (~$1.5B) goes to dedicated orchestration/agent infra vs $12.5B on raw model APIs — the harness tooling layer is still nascent relative to model spend.

$37B 2025 spend
Application layer iSoftware people actually use — copilots, agents, vertical apps. $19B / 51% of spend. Within its $8.4B horizontal slice, copilots dominate at $7.2B.$19B · 51%
Infrastructure layer$18B · 49%
↳ Foundation-model APIs$12.5B
↳ Training infrastructure$4B
↳ Data / storage / orchestration iDedicated AI infrastructure — including the orchestration / harness tooling you'd build — is only ~$1.5B of the $18B infra layer. Tiny vs the $12.5B spent on raw model APIs. Early = opportunity, but revenue is thin today.$1.5B

Adoption funnel — from interest to production iEach step filters the last. The big drop-offs are "prototype → true agent" and "true agent → production" — exactly the reliability/eval gap that defines the harness opportunity.

Enterprises building with gen-AI ~universal interest
Agents in production (builder-community view) 57%
"True agents", enterprise 16%
Production-live (strict screen) 5.2%
The 57% figure is from LangChain's builder survey (n=1,340; up from 51% the prior year; 67% at 10,000+ employee firms) — a self-selected pool of teams already building agents, so it runs high. The 5.2% is Cleanlab's stricter production screen. Both are directional, not census-grade. The gap between them is the reliability problem this whole market is trying to solve.

Market-sizing projections ($10.9B in 2026 → $182.9B by 2033 at 49.6% CAGR, Grand View Research) exist but are soft and diverge sharply by firm — treat as narrative, not planning. Menlo's spend data is the defensible anchor.
The mental map

02The harness stack — colored by who wins each layer

The entire space on one screen. Horizontal layers run from the model up to the agent; two rails (evaluation, governance) cut across everything. Color = the strategic verdict. Hover any layer's i…marker for who's winning it, why, and the provider-risk read. for the detail.

Agent interface / harness
The top-level "how you build & run an agent" surface
PROVIDER-OWNED iProvider risk: 5/5. This is exactly what OpenAI's AgentKit, Anthropic's Claude Agent SDK and Google's ADK want to own — it's the surface that keeps developers on their model. Don't build a standalone here.
Orchestration / frameworks
LangGraph · CrewAI · AutoGen · LlamaIndex · Temporal · Inngest
COMMODITIZING iProvider risk: 4/5. The open-source agent loop is a race to the bottom. Even LangChain fled up-stack into a proprietary platform (LangSmith). Durable execution (Temporal, $5B) is the exception — it's a broader infra category, not agent-specific.
Memory / context / state
Mem0 · Zep · Letta (MemGPT) · LlamaIndex
OPEN FRONTIER iProvider risk: 3/5. Bessemer's pick for the "next locus of differentiation as models commoditize." No category winner yet — genuinely open — but providers are adding native memory, so encroachment risk is real and monetization is unproven.
Retrieval / knowledge (RAG)
Vector DBs + retrieval logic
COMMODITIZED iThe vector database is rarely the bottleneck anymore — which chunks got retrieved is. Value is shifting from storage to retrieval quality, which blurs into the memory/context layer above.
Tools / connectivity (MCP)
MCP ecosystem · Composio · Arcade · LiteLLM (gateway)
STANDARDIZING iProvider risk: 4/5. MCP (authored & open-sourced by Anthropic) is now the de-facto standard — "publishing an MCP server is replacing writing a custom integration." Standardization compresses margins; LiteLLM has 418M monthly downloads on ~$15M raised — adoption ≫ monetization.
Runtime / sandboxing
E2B · Modal · Daytona · Vercel Sandbox · Fly
ENTRENCHED iProvider risk: 2/5. Real technical moat (isolation tech, cold-start latency) but a narrow, capital-intensive niche with entrenched players. Daytona ~90ms cold-start, E2B ~150ms (Firecracker microVMs), Modal sub-second + GPUs. Hard to enter late.
Models / inference / routing
The providers themselves + gateways (Portkey, OpenRouter)
PROVIDER TURF iThis is the providers' home turf. Routing/gateways are a thin, commoditizing wrapper on top. Not a pivot target.
▮ RAIL — Observability & Evaluation WEDGE ✦
Tracing is table-stakes; deep evaluation is the widest open gap in the whole stack. LangSmith · Braintrust · Arize · Langfuse (→ ClickHouse)
▮ RAIL — Governance & Security WEDGE ✦
The pilot-to-production gate for enterprises. Guardrails AI · Lakera · Protect AI (→ Palo Alto) · NeMo Guardrails
Provider-owned / commoditizing — avoid solo Consolidating / entrenched — hard Open / defensible — build here = provider-risk score
The landscape

03Who's who — the player map, by layer

The notable companies in each layer with verifiable funding, traction and positioning. Hover any row's i…marker for funding-round detail, benchmarks and source notes. for the deeper detail behind the cell. "P-risk" = how exposed the layer is to model-provider commoditization (5 = fully exposed).

LayerKey playersFunding / tractionPositioning readP-risk
Orchestration LangChain/LangGraph, LlamaIndex, CrewAI, AutoGen, Temporal, Inngest LangChain: $125M Series B
$1.25B val · Oct 2025 · IVP lediLangChain $125M Series B @ $1.25B (Oct 2025), led by IVP with Sequoia, Benchmark, CapitalG, Sapphire + strategics (ServiceNow, Workday, Cisco, Datadog, Databricks). Temporal: ~$430M cumulative, ~$5B val (Series D, Feb 2026), reported 380% YoY revenue growth, 20M+ monthly installs.LangChain blog + Fortune; presenc.ai. Verified 3-0. (A "35% of Fortune 500 use LangChain" claim was refuted & excluded.)
OSS core commoditizing; winners flee up-stack into proprietary lifecycle platforms ("agent engineering")
Eval / observability LangSmith, Braintrust, Arize/Phoenix, Langfuse, Galileo, Helicone, W&B Weave, Fiddler Braintrust: ~$800M val
~$120M raised (Series B)iBraintrust: ~$120M cumulative, ~$800M valuation (Series B); differentiates on CI/CD gating (auto-block merges on eval fail). Langfuse acquired by ClickHouse (Jan 2026), part of a $400M Series D — consolidation underway. Pricing spans free → ~$249/mo (Braintrust), $39/user/mo (LangSmith).braintrust.dev; presenc.ai. Verified.
Observability = table-stakes (89–94%); the eval half is under-served. Differentiate on depth + CI/CD gating
Runtime / sandbox E2B, Modal, Daytona, Vercel Sandbox, Fly Cold-start: Daytona ~90ms
E2B ~150ms · Modal sub-seciCold-start latency: Daytona ~90ms (27ms optimized) < E2B ~150ms < Modal sub-second. Isolation tech: E2B = Firecracker microVMs (dedicated kernel); Modal = gVisor + the only one with GPU sandboxes (T4–H100); Daytona = containers. CPU price: Daytona/E2B $0.0504/vCPU-hr, Modal ~$0.1419 (3× premium).startuphub.ai benchmark (blog, 2026).
Genuine technical moat but narrow & capital-heavy; hard to enter late
Memory / context Mem0, Zep, Letta (MemGPT), LlamaIndex Early-stage; no clear leaderiBessemer's 2026 roadmap names memory/context the emerging locus of differentiation "as models become commoditized." Dense emerging ecosystem, but no breakout funding leader identified in verified sources — the frontier is genuinely open (and unproven).Bessemer AI Infrastructure Roadmap 2026 (primary). Hot, open, no winner — but provider-encroachment risk (native memory) & unproven monetization
Tools / MCP MCP ecosystem, Composio, Arcade, LiteLLM, Portkey LiteLLM: 418M/mo downloads
on ~$15M raisediLiteLLM (BerriAI): ~418M monthly PyPI downloads on ~$15M cumulative funding — "the most-deployed agent-infra component," but monetization badly lags adoption. The cautionary tale of this layer: ubiquity ≠ revenue.presenc.ai (blog, 2026).
MCP standardized fast (Anthropic-authored, OSS); margins compressed, monetization unproven
Guardrails / security Guardrails AI, Lakera, Protect AI, NeMo Guardrails (Nvidia) Protect AI → Palo Alto (~$500M)
MCP-security only ~$40M fundediConsolidation into security incumbents: Palo Alto acquired Protect AI (~$500M). Top-10 agentic-AI-security startups raised ~$3.6B combined, but concentrated (Cyera $1.7B+, Saviynt ~$1.0B). MCP-specific security is still tiny — ~$40M across 4 startups (Operant AI, Runlayer, Helmet, Manufact): nascent, underfunded, wide open.softwarestrategiesblog.com (blog, Mar 2026).
Broad security consolidating into incumbents; but agent-native security still forming & open
Pattern across the table: every layer is either (a) fronted by a well-funded leader already up-stack, (b) consolidating into incumbents (observability → ClickHouse; security → Palo Alto), or (c) genuinely early with no winner (memory, agent-native security). Bucket (c) is where a new entrant has room. Note: granular "X% of agent funding went to layer Y" claims circulate widely but were refuted in fact-checking — excluded here.
What changed

04How the ground shifted — 2023 → 2026

The market's shape changed fast. Two forces dominate the timeline: standardization (MCP) collapsing the integration layer, and the model providers moving up-stack to bundle the harness. Hover each node for detail.

2023
Code Interpreter era
OpenAI function-calling + code execution. E2B is "first sandbox built around the Code Interpreter pattern" — the runtime layer is born.
Nov 2024
Anthropic ships MCP
Model Context Protocol released & open-sourced. Over ~18 months it becomes the de-facto standard for tool connectivity — collapsing the integration layer.
H1 2025
Framework proliferation
LangGraph, CrewAI, AutoGen, LlamaIndex multiply. Orchestration becomes abundant → commoditizing.
Aug 2025
Category recognized
CB Insights publishes "The AI Agent Tech Stack" — 135+ startups across 17 sub-markets. The harness is now a named analyst category.
Oct 2025
OpenAI AgentKit ✦
DevDay: Agent Builder + ChatKit + Connector Registry + Evals + Guardrails. One provider bundles 5 harness layers. The pivotal platform-risk event.
Oct 2025
LangChain → unicorn
$125M Series B @ $1.25B, led by IVP. Repositions around "agent engineering" & proprietary LangSmith — the strongest independent flees up-stack.
Jan 2026
Langfuse → ClickHouse
Observability consolidation begins. Part of ClickHouse's $400M Series D.
2026
Governance becomes the gate
Enterprise conversations shift: governance moves from near-absent (2025) to the factor separating pilots from production. Protect AI → Palo Alto (~$500M).
The net effect: the commodity line has crept upward and outward. Layers that were venture-backed startup categories in early 2025 (frameworks, connectivity, basic evals) are, by 2026, either provider-bundled or consolidating. The value migrated to the two rails — evaluation depth and governance — plus vertical and vendor-neutral positioning.
Where the pain actually is

05Reliability is the #1 problem — and cost is fading

Ranked barriers to shipping agents in production. The strategic signal: "help me make it work & be safe" is a growing pain; "help me spend less" is shrinking as model prices fall. Build toward the growing pain.

Quality / reliability
32% · #1
Security (enterprise) i24.9% — the #2 barrier specifically for organizations with 2,000+ employees. Prompt injection, data exfiltration, and agent permissioning are enterprise-gating. Data privacy/compliance is repeatedly cited as the primary restraint in regulated industries.LangChain State of Agent Engineering 2025. Verified 3-0.
24.9%
Latency
20% · #2 overall
Cost ↓ falling iCost concern is declining year-over-year as model prices drop. Strategically important: a product whose whole pitch is "spend less on tokens" is attacking a shrinking pain. Reliability & safety are the growing ones.LangChain 2025; Cleanlab 2025.
declining

The same pain, different by scale iA crucial GTM nuance: which pain dominates depends on the buyer's scale. This changes who you sell to and what you lead with.

Buyer profileDominant painWhat they'll pay forWillingness to pay
Small / early deploymentsCostToken efficiency, cheap toolingLow ($ / bottoms-up)
High-traffic agentsLatency + reliabilityPerformance, eval, uptimeMedium
Large / regulated enterprisesSecurity + governanceAudit, compliance, guardrails, HITLHigh ($$$ ACV)
Startup vs enterprise, distilled: startups feel velocity pain (ship fast, prototype→prod gap, cost) — adopt bottoms-up but pay little. Enterprises feel trust pain (reliability, security, governance) — pay real ACV but move slowly and demand self-hosting & audit. The money is in enterprise trust; the volume is in startup velocity. Sources: LangChain (n=1,340) & Cleanlab (n=1,837), 2025 — vendor surveys, directional.
Solved vs unsolved

06Teams can see their agents — they can't judge them

The single widest gap in the stack. Observability is near-universal, but evaluation is not — teams have dashboards full of traces and no reliable way to know if the output was actually right. That gap is the clearest wedge in the whole analysis.

Have observability
89–94%
Run offline evals
52.4%
Run online evals
37.3%
Run NO evals at all i22.8% of production teams run no evaluation whatsoever — shipping non-deterministic systems blind. Combined with the ~78% of failures that are silent, this is the reliability crisis in one number.LangChain 2025 (verified 3-0); "78% invisible failures" from Bessemer citing the WildChat academic study.
22.8%
Satisfied with tooling
< 1 in 3
Observability is also the #1 forward investment priority — 62% plan improvements. Fewer than 1 in 3 teams are satisfied with their current observability/guardrails. And ~78% of AI failures are "invisible" — plausible-but-wrong outputs no one catches — which is precisely what deeper eval + silent-failure detection would surface.

Solved — don't build here

  • The agent loop / basic orchestration — free & abundant
  • Tool connectivity — MCP standardized it in ~18 months
  • Tracing plumbing — near-universal, table-stakes
  • Model routing — LiteLLM et al. ubiquitous (418M downloads/mo)
  • Retrieval mechanics — vector DBs commoditized

Still painful — build here

  • Deep evaluation — the ~half of teams flying blind
  • Silent-failure detection — ~78% of failures are invisible
  • Memory / long-horizon context — no winner yet
  • Agent-native security — prompt injection, permissioning
  • Enterprise governance — the pilot → production gate
Platform risk

07What the model providers will absorb

The existential question for any harness pivot. The providers have three unfair advantages — distribution, the model, and the ability to bundle for free — and they're already using them. The table shows what each is shipping, and the verdict below sorts the stack into "absorbed" vs "stays independent".

ProviderAgent toolkitWhat it bundlesEnterprise API share
OpenAIAgentKit (Oct 2025)Agent Builder (orchestration), ChatKit (UI), Connector Registry (MCP), Evals, Guardrails iOne release bundling 5 previously-independent harness layers. Notably, Evals includes third-party-model support — directly targeting the vendor-neutral eval niche — and Guardrails is open-source. So not fully provider-locked, but the bundling is the platform-risk signal.openai.com/index/introducing-agentkit (primary). Verified 3-0.27% (↓ from 50% in '23)
AnthropicClaude Agent SDK + MCPAuthored & open-sourced MCP (the connectivity standard); agent SDK, tool use, memory iAnthropic owns the connectivity standard itself. Enterprise API share climbed to 40% (up from 24% in 2024, 12% in 2023), and 54% in coding specifically — concentration that raises platform dependence for anyone building on top.Menlo Ventures 2025. Verified supporting.40% · 54% coding
GoogleADK (Agent Dev Kit)Orchestration, tool use, Gemini-native memory & code execution21%

⤵ Gets absorbed — high platform risk

  • ✕ Orchestration frameworks / the agent loop
  • ✕ Basic tool connectivity (MCP is provider-blessed)
  • ✕ Single-vendor evals & basic guardrails
  • ✕ Embeddable chat UI
  • ✕ Basic memory (going native)

⤴ Stays independent — structural moat

  • Cross-vendor / neutral tooling — a conflict of interest they can't fully lean into
  • Deep, opinionated evaluation beyond a bundled checkbox
  • Enterprise governance / compliance — not their DNA
  • Vertical depth — domain-specific agents & their harness
  • Multi-provider orchestration at enterprise scale
The structural insight: a model provider's eval or gateway that "also supports competitors" is a conflict of interest they'll never fully commit to — so neutrality is a moat they cannot copy without undermining their own model business. That's why AgentKit's third-party-model support is real but half-hearted, and why "teams deliberately choosing neutral tooling to avoid lock-in" is a durable behavior to build on.
The wedges left

08Defensibility scorecard & map

Each candidate wedge scored on the factors that decide a pivot. Hover a row's i…marker for the reasoning behind the tier. for reasoning. Then the 2-D map places them by "is the pain growing?" (up) vs "can a provider own it?" (right = no = defensible).

WedgePain & growthProvider-proof?Moat typeSales / ACVTier
Enterprise governance, security & complianceiReal, enterprise-gating pain that's growing (governance became the pilot→prod gate in 2026). Providers optimize for developers, not the CISO/auditor/regulator. Moat = certifications, audit trails, trust relationships. Best risk-adjusted wedge, especially with a regulated-vertical focus. High & growing Strong ✓✓Trust, compliance, certsSlow · High ACV Tier 1
Deep evaluation / reliabilityiWidest solved-vs-unsolved gap (89% observe, ~half evaluate, <⅓ satisfied). Providers bundle basic evals, so you must go deeper: domain-specific scoring, silent-failure detection, regression gating. Defensible if you own proprietary eval data/methods and stay cross-vendor. Highest gap Strong ✓✓Proprietary eval data + neutralityMid · Medium–High Tier 1
Vendor-neutral, cross-model toolingiA structural moat a provider can't copy without cannibalizing itself. But "neutral" alone isn't a business — it works as a property of the governance or eval wedge, not a standalone product. Enabler Structural ✓✓✓Conflict-of-interest moatVaries Tier 1*
Memory / context layeriAnalyst-favored (Bessemer), genuinely open — but high provider-encroachment risk (native memory) and no proven monetization model yet. High upside, high uncertainty. Emerging Medium ~Technical depth (unproven)Unclear Tier 2
Vertical agent + its harnessiThe application layer (vertical agents) is where most revenue actually is today. A hybrid — a vertical agent that productizes its own eval/governance harness — can be more defensible than either alone: revenue now + infra moat later. Revenue now Medium ~Domain data + workflow lock-inVaries · Direct rev Tier 2
Frameworks · MCP hubs · runtime · basic observabilityiCommoditized, provider-owned, margin-compressed, or consolidating into incumbents (Langfuse→ClickHouse, Protect AI→Palo Alto). Do not enter as a standalone business. Solved / commoditized Weak ✕—— Avoid solo

The 2-D defensibility map

The scorecard above carries the full ranking — the 2-D map needs a wider screen.
▲ pain real & growing
◀ provider can own it (avoid)
hard for provider to own ▶
Enterprise governance & security
Deep evaluation / reliability
Vendor-neutral tooling
Memory / context
Vertical agent + harness
Orchestration frameworks
MCP / connectivity hubs
Basic observability
Top-right is where you want to be: pain that's real and growing, in a place a model provider can't profitably follow. The two Tier-1 wedges (governance, deep eval) sit there; the commodity layers sit bottom-left.
The verdict

09If you pivot — three shapes, one rule

Pivot into reliability/eval or governance, not framework/connectivity. Which of the three shapes fits is decided entirely by Nometria's unfair advantage.

AEnterprise governance & reliability platform

Sell trust to platform teams & CISOs at mid/large & regulated firms: eval, silent-failure detection, guardrails, audit, human-in-the-loop, compliance — cross-vendor by design.

Fits if: you have enterprise GTM or regulated-domain relationships. Highest ACV, most defensible, slowest sales.
BVendor-neutral evaluation infrastructure

Attack the eval gap head-on — the ~half of teams flying blind. Go deeper than bundled evals: domain scoring, regression gating, reliability SLAs.

Fits if: your strength is data / ML / eval. Defensible via proprietary methods + neutrality.
CVertical agent that productizes its harness

Pick one high-value domain, build the agent and the eval/governance layer it needs, then consider spinning the harness out.

Fits if: you have deep domain expertise in a vertical. Revenue now + a moat later.

Which shape fits Nometria? — decision guide

If…
your edge is enterprise relationships / a regulated domain (finance, health, legal)
→ Shape A
If…
your edge is data / ML / evaluation engineering depth
→ Shape B
If…
your edge is deep expertise in one vertical + existing users/workflow
→ Shape C
The one rule: design around the assumption that model providers keep absorbing the commodity middle. Your moat must live where they can't profitably follow — neutrality, enterprise trust, vertical depth, evaluation rigor. Anything a future AgentKit release could bundle for free is not a business.
The map can't pick A / B / C for you — that's set by what Nometria is today and where the team is strongest. Tell me that and I'll pressure-test the specific wedge, size it, name the nearest incumbent to beat, and sketch a concrete v1.
Sources — Primary: Menlo Ventures, State of Generative AI in the Enterprise 2025 PRIMARY · Bessemer, AI Infrastructure Roadmap 2026 PRIMARY · LangChain, State of Agent Engineering & Series B PRIMARY · OpenAI, AgentKit PRIMARY · Cleanlab, AI Agents in Production 2025 PRIMARY. Secondary: CB Insights, Grand View Research, O'Reilly, Braintrust, StartupHub, SoftwareStrategies, presenc.ai SECONDARY / BLOG.

Methodology & caveats: 24 sources fetched → 117 claims extracted → 25 adversarially fact-checked (22 confirmed, 3 refuted & excluded). Adoption/pain figures come from vendor-run surveys of self-selected agent builders — directional, not census-grade (explains the 5.2%–57% production-adoption spread). Market-sizing projections diverge sharply by firm; Menlo's spend data is the defensible anchor. The "78% invisible failures" stat is an academic study Bessemer cites, not their own estimate. Excluded refuted claims: "$4.7B raised across 59 agentic deals," "vertical agents = 55.7% of capital," "35% of Fortune 500 use LangChain." Fast-moving space (AgentKit & the LangChain raise are Oct 2025) — revalidate quarterly.