Open source has already commoditized the primitives — PII detection, prompt-injection classifiers, policy engines, tracing, red-team scanners. It has not built the thing that's actually defensible: a unified, vendor-neutral control plane that ties runtime enforcement to deep evaluation, immutable audit, and compliance mapping. Reuse aggressively below the value line; spend your engineering above it.
Actively maintained, permissively licensed primitives. Adopt directly and put your UX/control plane on top.
Useful to study or borrow patterns, but archived, stale, or now owned by a competitor — building your foundation on them is a trap.
No OSS project does these well. This is the product — the reason to exist.
Grouped by what they do (which maps to your pillars). License, health and verdict are the columns that decide build-vs-reuse. Watch the license colors: several popular "open" safety models are not permissively licensed.
| Project | Maintainer | What it does | License | Health | Pillar | Verdict |
|---|---|---|---|---|---|---|
| Guardrail orchestration frameworks (run & compose checks) → Pillar 3 | ||||||
| NeMo Guardrails | NVIDIA | Programmable rails (Colang) — topical, jailbreak, dialogue, fact-check; composes checks | Apache-2.0 | Active | 3 | Reuse |
| Guardrails AI | Guardrails AI | I/O validation, structured output, "Guardrails Hub" of community validators + auto-correction | Apache-2.0iCore is Apache-2.0, but individual Hub validators can carry their own licenses — check per-validator before shipping. | Active | 3 | Reuse |
| LLM Guard | Protect AI (Palo Alto) | 15 input + 20 output scanners (injection, PII, toxicity, secrets, relevance…) | MIT | Archived Jul '26 | 3 | Reference |
| OpenAI Guardrails | OpenAI (AgentKit) | Provider-blessed guardrails config; open-source but OpenAI-centric | MIT | Active | 3 | Reference |
| Prompt-injection / jailbreak & safety classifiers → Pillar 3 | ||||||
| Rebuff | Protect AI | Multi-layer prompt-injection detection (heuristics + LLM + canary tokens) | Apache-2.0 | Stale | 3 | Reference |
| Vigil | community | Prompt-injection / jailbreak scanner & detection library | Apache-2.0 | Low activity | 3 | Reference |
| Llama Guard 4 / Prompt Guard 2 | Meta | Safety-classifier models for input/output moderation & injection | Llama CommunityiNOT OSI-approved. The Llama Community License adds an Acceptable-Use Policy and a >700M-MAU clause. Usable for most startups, but it's a restricted license — don't treat it as permissive, and factor it into a commercial product's legal review. | Active | 3 | Reuse ⚠ |
| Granite Guardian | IBM | Safety/risk classifier models (harm, jailbreak, RAG hallucination checks) | Apache-2.0 | Active | 3/4 | Reuse ★iThe cleanest license in the safety-classifier group — true Apache-2.0 weights. Prefer this over Llama Guard / ShieldGemma when you need permissive terms for a commercial product. |
| ShieldGemma | Safety-classifier models on content-policy categories | Gemma License | Active | 3 | Reuse ⚠ | |
| PII / data protection → Pillar 3 | ||||||
| Presidio | Microsoft | PII detection, anonymization & de-identification SDK (text + images); customizable recognizers | MIT | Active ★ | 3 | Reuse ★iThe de-facto OSS standard for PII. Mature, permissive, extensible. Rebuilding PII detection from scratch would be pure duplicated work — wrap Presidio. |
| Red-teaming, eval & testing → Pillar 4 | ||||||
| Garak | NVIDIA | LLM vulnerability scanner — probes injection, jailbreak, data leakage, toxicity | Apache-2.0 | Active | 4 | Reuse |
| PyRIT | Microsoft | Python Risk Identification Toolkit — automated, orchestrated red-teaming of GenAI | MIT | Active | 4 | Reuse |
| promptfoo | promptfoo | Eval + red-team harness with CI integration & regression gating | MIT | Active ★ | 4 | Reuse ★iStrong fit for the Phase-0 "eval gates in CI" MVP feature. Reuse as the eval-runner substrate; your value-add is domain-specific scorers + silent-failure detection + the compliance tie-in. |
| Giskard | Giskard | LLM testing & scan — vulnerabilities, bias, hallucination; framework-agnostic | Apache-2.0 | Active | 4 | Reuse |
| Policy & authorization engines → Pillar 2 / 6 | ||||||
| Open Policy Agent (OPA) | CNCF | General policy engine (Rego); decouple policy-as-code from app; agent tool authz | Apache-2.0 | Active ★ | 2/6 | Reuse ★iCNCF-graduated, battle-tested. Your "policy-as-code engine" pillar should be OPA/Rego (or Cedar) under the hood, with an agent-native policy authoring layer on top — not a bespoke engine. |
| Cedar | AWS | Fine-grained authorization language/engine; formally verified; sub-ms decisions | Apache-2.0 | Active | 2 | Reuse |
| Casbin | community | ACL / RBAC / ABAC authorization across many languages | Apache-2.0 | Active | 2 | Reuse |
| Observability & audit → Pillar 5 | ||||||
| OpenTelemetry (+ OpenLLMetry) | CNCF / Traceloop | Distributed tracing/metrics/logs standard; LLM semantic conventions | Apache-2.0 | Active ★ | 5 | Reuse ★ |
| Langfuse | Langfuse (ClickHouse) | LLM observability — traces, spans, eval, cost; self-hostable | MITiCore is MIT and self-hostable, but the company was acquired by ClickHouse (Jan 2026). Fine as a component; be aware some features sit behind a commercial/EE tier. | Acquired | 5 | Reuse/Ref |
| Helicone | Helicone | Proxy for request logging, cost tracking, rate limiting | Apache-2.0 | Active | 5 | Reuse |
| Agent / MCP-native security → Pillar 1 / 3 | ||||||
| Invariant Guardrails / mcp-scan | Snyk (acq.) | Agent/MCP-native guardrails; scans MCP servers for tool-poisoning & injection | Apache-2.0 | Snyk-owned | 1/3 | Reference |
| systemprompt-core | systempromptio | MCP governance runtime — authn/authz, rate-limit, logging (Rust) | BSL-1.1iBusiness Source License — "source-available," NOT open source. Converts to Apache after a delay, but has usage restrictions today. Don't build a commercial product on a BSL dependency without legal review. | Active | 1/6 | Reference ⚠ |
| Microsoft Agent Governance Toolkit | Microsoft | Runtime security across 15+ agent frameworks; OWASP Agentic Top-10 coverage | MIT | Active | 1/3 | Reference |
| Standards & taxonomies (not code — adopt as your control language) | ||||||
| OWASP LLM / Agentic Top 10 | OWASP | Canonical risk taxonomy for LLM & agent threats | Open | Active | 3/6 | Adopt |
| MITRE ATLAS | MITRE | Adversarial ML threat matrix — map your detections to it | Open | Active | 3 | Adopt |
How much of each of your six pillars you can assemble from open source today. The pattern is the whole strategy: OSS covers the security & testing primitives well, thins out on audit & identity, and effectively doesn't exist for compliance — which is exactly where your defensibility lives.
A clean, Apache/MIT-only foundation for the Phase-0 MVP (inline guardrails + trace + eval-gating), with nothing archived, competitor-owned, or license-restricted on the critical path.
| Layer | Pick | License | Why this one |
|---|---|---|---|
| PII detection/redaction | Presidio | MIT | De-facto standard, extensible, mature |
| Safety classifier | Granite Guardian | Apache-2.0 | Cleanest license of the classifier set |
| Guardrail orchestration | NeMo Guardrails or Guardrails AI | Apache-2.0 | Compose checks; don't hand-roll the runner |
| Eval + CI gating | promptfoo | MIT | Purpose-built for regression gating in CI |
| Red-teaming | Garak + PyRIT | Apache / MIT | Scanner + orchestrated adversarial suites |
| Authorization | OPA (Rego) or Cedar | Apache-2.0 | Proven policy engines; sub-ms decisions |
| Tracing / audit spine | OpenTelemetry + OpenLLMetry | Apache-2.0 | Industry standard; feeds your evidence layer |
| Risk taxonomy | OWASP Agentic Top 10 + MITRE ATLAS | Open | Map detections to a language buyers trust |