Open-Source Guardrails — Build-vs-Reuse Catalog

So Nometria doesn't rebuild commodity plumbing. Every meaningful open-source guardrail / LLM-security / red-team / policy project, with its license, health, what it covers, and a verdict — reuse it, reference it, or build past it. Mapped to the six roadmap pillars. Hover any i…marker for detail, license nuance, or source..
Scanned: ~25 OSS projects across 7 categories Grounded: project repos + 2026 comparison sources Watch: 2 key projects archived / acquired this year
The bottom line first

01Reuse the plumbing, build the differentiation

Open source has already commoditized the primitives — PII detection, prompt-injection classifiers, policy engines, tracing, red-team scanners. It has not built the thing that's actually defensible: a unified, vendor-neutral control plane that ties runtime enforcement to deep evaluation, immutable audit, and compliance mapping. Reuse aggressively below the value line; spend your engineering above it.

One-line rule: if an Apache/MIT project already does a primitive well and is actively maintained, wrap it — don't rebuild it. Your product is the integration + eval depth + compliance mapping + neutrality, none of which any single OSS project provides.

✓ Reuse — wrap, don't rebuild

Actively maintained, permissively licensed primitives. Adopt directly and put your UX/control plane on top.

  • Presidio — PII detection/redaction (MIT)
  • OPA · Cedar · Casbin — policy / authorization
  • OpenTelemetry — trace/audit backbone
  • Garak · PyRIT · promptfoo · Giskard — red-team & eval engines
  • NeMo Guardrails · Guardrails AI — guardrail orchestration frameworks
  • Granite Guardian — safety classifier (clean Apache license)

◑ Reference only — don't depend on

Useful to study or borrow patterns, but archived, stale, or now owned by a competitor — building your foundation on them is a trap.

  • LLM Guard — archived Jul 2026 (Protect AI → Palo Alto) iExcellent scanner design (15 input + 20 output scanners, MIT), but the repo and its HF models were archived on 9 Jul 2026 — no longer maintained. Mine it for scanner ideas; don't build on a dead dependency.
  • Rebuff — Protect AI, effectively stale
  • Invariant guardrails · mcp-scan — acquired by Snyk iThe most agent/MCP-native OSS guardrails — now owned by Snyk (an incumbent building the same category). Fine to use mcp-scan as a tool; risky to build your differentiation on a competitor's roadmap.
  • Llama Guard / Prompt Guard — license-restricted (see §02)

✦ Build — your white space

No OSS project does these well. This is the product — the reason to exist.

  • Unified control plane across all primitives
  • Deep domain eval + silent-failure detection
  • Compliance mapping (EU AI Act / NIST / ISO 42001)
  • Immutable audit + auditor-ready evidence
  • Cross-vendor neutrality + agent registry/discovery
The landscape

02The catalog — every project, by category

Grouped by what they do (which maps to your pillars). License, health and verdict are the columns that decide build-vs-reuse. Watch the license colors: several popular "open" safety models are not permissively licensed.

ProjectMaintainerWhat it doesLicenseHealthPillarVerdict
Guardrail orchestration frameworks (run & compose checks) → Pillar 3
NeMo GuardrailsNVIDIAProgrammable rails (Colang) — topical, jailbreak, dialogue, fact-check; composes checksApache-2.0Active3Reuse
Guardrails AIGuardrails AII/O validation, structured output, "Guardrails Hub" of community validators + auto-correctionApache-2.0iCore is Apache-2.0, but individual Hub validators can carry their own licenses — check per-validator before shipping.Active3Reuse
LLM GuardProtect AI (Palo Alto)15 input + 20 output scanners (injection, PII, toxicity, secrets, relevance…)MITArchived Jul '263Reference
OpenAI GuardrailsOpenAI (AgentKit)Provider-blessed guardrails config; open-source but OpenAI-centricMITActive3Reference
Prompt-injection / jailbreak & safety classifiers → Pillar 3
RebuffProtect AIMulti-layer prompt-injection detection (heuristics + LLM + canary tokens)Apache-2.0Stale3Reference
VigilcommunityPrompt-injection / jailbreak scanner & detection libraryApache-2.0Low activity3Reference
Llama Guard 4 / Prompt Guard 2MetaSafety-classifier models for input/output moderation & injectionLlama CommunityiNOT OSI-approved. The Llama Community License adds an Acceptable-Use Policy and a >700M-MAU clause. Usable for most startups, but it's a restricted license — don't treat it as permissive, and factor it into a commercial product's legal review.Active3Reuse ⚠
Granite GuardianIBMSafety/risk classifier models (harm, jailbreak, RAG hallucination checks)Apache-2.0Active3/4Reuse ★iThe cleanest license in the safety-classifier group — true Apache-2.0 weights. Prefer this over Llama Guard / ShieldGemma when you need permissive terms for a commercial product.
ShieldGemmaGoogleSafety-classifier models on content-policy categoriesGemma LicenseActive3Reuse ⚠
PII / data protection → Pillar 3
PresidioMicrosoftPII detection, anonymization & de-identification SDK (text + images); customizable recognizersMITActive ★3Reuse ★iThe de-facto OSS standard for PII. Mature, permissive, extensible. Rebuilding PII detection from scratch would be pure duplicated work — wrap Presidio.
Red-teaming, eval & testing → Pillar 4
GarakNVIDIALLM vulnerability scanner — probes injection, jailbreak, data leakage, toxicityApache-2.0Active4Reuse
PyRITMicrosoftPython Risk Identification Toolkit — automated, orchestrated red-teaming of GenAIMITActive4Reuse
promptfoopromptfooEval + red-team harness with CI integration & regression gatingMITActive ★4Reuse ★iStrong fit for the Phase-0 "eval gates in CI" MVP feature. Reuse as the eval-runner substrate; your value-add is domain-specific scorers + silent-failure detection + the compliance tie-in.
GiskardGiskardLLM testing & scan — vulnerabilities, bias, hallucination; framework-agnosticApache-2.0Active4Reuse
Policy & authorization engines → Pillar 2 / 6
Open Policy Agent (OPA)CNCFGeneral policy engine (Rego); decouple policy-as-code from app; agent tool authzApache-2.0Active ★2/6Reuse ★iCNCF-graduated, battle-tested. Your "policy-as-code engine" pillar should be OPA/Rego (or Cedar) under the hood, with an agent-native policy authoring layer on top — not a bespoke engine.
CedarAWSFine-grained authorization language/engine; formally verified; sub-ms decisionsApache-2.0Active2Reuse
CasbincommunityACL / RBAC / ABAC authorization across many languagesApache-2.0Active2Reuse
Observability & audit → Pillar 5
OpenTelemetry (+ OpenLLMetry)CNCF / TraceloopDistributed tracing/metrics/logs standard; LLM semantic conventionsApache-2.0Active ★5Reuse ★
LangfuseLangfuse (ClickHouse)LLM observability — traces, spans, eval, cost; self-hostableMITiCore is MIT and self-hostable, but the company was acquired by ClickHouse (Jan 2026). Fine as a component; be aware some features sit behind a commercial/EE tier.Acquired5Reuse/Ref
HeliconeHeliconeProxy for request logging, cost tracking, rate limitingApache-2.0Active5Reuse
Agent / MCP-native security → Pillar 1 / 3
Invariant Guardrails / mcp-scanSnyk (acq.)Agent/MCP-native guardrails; scans MCP servers for tool-poisoning & injectionApache-2.0Snyk-owned1/3Reference
systemprompt-coresystempromptioMCP governance runtime — authn/authz, rate-limit, logging (Rust)BSL-1.1iBusiness Source License — "source-available," NOT open source. Converts to Apache after a delay, but has usage restrictions today. Don't build a commercial product on a BSL dependency without legal review.Active1/6Reference ⚠
Microsoft Agent Governance ToolkitMicrosoftRuntime security across 15+ agent frameworks; OWASP Agentic Top-10 coverageMITActive1/3Reference
Standards & taxonomies (not code — adopt as your control language)
OWASP LLM / Agentic Top 10OWASPCanonical risk taxonomy for LLM & agent threatsOpenActive3/6Adopt
MITRE ATLASMITREAdversarial ML threat matrix — map your detections to itOpenActive3Adopt
Permissive Apache / MIT — safe to build on Caution Gemma / community — restricted Restricted Llama / BSL — legal review needed ★ = best-in-class pick
Where OSS helps — and where it doesn't

03Coverage map — OSS availability per roadmap pillar

How much of each of your six pillars you can assemble from open source today. The pattern is the whole strategy: OSS covers the security & testing primitives well, thins out on audit & identity, and effectively doesn't exist for compliance — which is exactly where your defensibility lives.

3 · Runtime Guardrails & Securityinjection, PII, toxicity, tool-call
Well covered
4 · Evaluation & Reliabilityred-team, eval, regression
Good primitives
2 · Identity, Access & Authorizationpolicy engines, RBAC/ABAC
Engines yes, agent-native no
5 · Audit, Observability & Traceabilitytracing yes; immutable/evidence no
Partial
1 · Discovery & Agent RegistryMCP scan exists; registry doesn't
Thin / competitor-owned
6 · Policy & Compliance MappingEU AI Act / NIST / ISO 42001
Essentially none
Read top-to-bottom: the more table-stakes and commoditized a pillar, the more OSS you get for free (guardrails, eval). The more it's your moat — compliance mapping, immutable audit, agent-native identity/registry, cross-vendor neutrality — the less OSS exists, because it requires product integration and domain work, not a library. Reuse where the bars are green; build where they're red.
Your engineering budget

04What to actually build (the non-duplicated work)

Assemble from OSS (weeks, not quarters)

  1. Runtime checks — wrap Presidio (PII) + Granite Guardian / Llama Guard (safety) + NeMo/Guardrails-AI (orchestration) behind one policy interface
  2. Eval & red-team — promptfoo for CI gating; Garak + PyRIT for adversarial suites
  3. Authorization — OPA (Rego) or Cedar for tool-scoped least-privilege
  4. Tracing — OpenTelemetry + OpenLLMetry as the span/trace backbone
  5. MCP hygiene — mcp-scan as a scanner (tool only, not a dependency)

Build yourself (the moat — no OSS does it)

  1. The unifying control plane — one policy model, one dashboard, one audit trail across every wrapped primitive & every model/cloud
  2. Deep evaluation + silent-failure detection — domain-specific scorers beyond generic red-team; the reliability layer that separates you from pure-security tools
  3. Compliance mapping engine — controls → EU AI Act / NIST AI RMF / ISO 42001 / SOC 2, with continuous evidence
  4. Immutable, tamper-evident audit + one-click auditor evidence — OTel gives spans; the evidentiary layer is yours
  5. Agent registry + discovery + cross-vendor neutrality — the control plane no model provider or single-suite vendor will build
Why this split is the point: the OSS column is where competitors also get their primitives — matching them there is table-stakes, not differentiation. The build column is what a model provider (conflict of interest), a GRC incumbent (no runtime/agent depth), or a point-security tool (no eval/compliance) structurally won't produce. Spend accordingly: ~20% of engineering integrating OSS, ~80% on the build column.
Concrete

05Recommended starter stack (all permissive)

A clean, Apache/MIT-only foundation for the Phase-0 MVP (inline guardrails + trace + eval-gating), with nothing archived, competitor-owned, or license-restricted on the critical path.

LayerPickLicenseWhy this one
PII detection/redactionPresidioMITDe-facto standard, extensible, mature
Safety classifierGranite GuardianApache-2.0Cleanest license of the classifier set
Guardrail orchestrationNeMo Guardrails or Guardrails AIApache-2.0Compose checks; don't hand-roll the runner
Eval + CI gatingpromptfooMITPurpose-built for regression gating in CI
Red-teamingGarak + PyRITApache / MITScanner + orchestrated adversarial suites
AuthorizationOPA (Rego) or CedarApache-2.0Proven policy engines; sub-ms decisions
Tracing / audit spineOpenTelemetry + OpenLLMetryApache-2.0Industry standard; feeds your evidence layer
Risk taxonomyOWASP Agentic Top 10 + MITRE ATLASOpenMap detections to a language buyers trust
Deliberately kept off the critical path: LLM Guard (archived), Rebuff/Vigil (stale), Invariant/mcp-scan (Snyk-owned — use as a tool, not a dependency), Llama Guard & ShieldGemma (license-restricted — swap in only if you accept the terms), systemprompt-core (BSL). Revisit these as references, not foundations.
Sources — Project repos & docs: protectai/llm-guard (archived notice, Jul 2026), protectai/rebuff, guardrails-ai, NVIDIA NeMo Guardrails & Garak, Microsoft Presidio & PyRIT, promptfoo, Giskard, Open Policy Agent, AWS Cedar, OpenTelemetry, IBM Granite Guardian, Meta Llama Guard / Prompt Guard, Google ShieldGemma. Landscape & comparisons: Giskard "Best AI guardrail tools 2026," awesome-ai-agent-governance (systempromptio), Snyk "acquires Invariant Labs" (agentic AI security), 2026 red-team & guardrail buyer guides (Galileo, appsecsanta, arXiv 2410.16527 OSS scanner comparison; ICLR 2026 safety-guard benchmark).

Caveats: license and maintenance status change quickly — verify each project's LICENSE file and last-commit date before committing it to your build. "Health" labels are as of this scan (Aug 2026). Model licenses (Llama Community, Gemma) and source-available licenses (BSL-1.1) are not OSI-approved open source and need legal review for a commercial product. Pillar numbers map to the Agent Governance Roadmap artifact.