Metadata-Version: 2.4
Name: ckg-nvidia-nemoclaw
Version: 0.13.1
Summary: Auditable knowledge graph for NVIDIA NemoClaw with cryptographic source traceability — 3.8× F1 over RAG, 11× fewer tokens, MCP-native, deterministic agent answers
Project-URL: Homepage, https://graphifymd.com
Project-URL: Documentation, https://yarmoluk.github.io/ckg-nvidia-nemoclaw
Project-URL: Benchmark, https://github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf
Author-email: Daniel Yarmoluk <daniel.yarmoluk@gmail.com>
License: MIT
License-File: LICENSE
License-File: LICENSE-CODE
License-File: LICENSE-DATA
Keywords: ai-agents,ckg,hermes,inference-routing,knowledge-graph,mcp,nemoclaw,network-policy,nvidia,openclaw,openshell,sandbox
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Security
Requires-Python: >=3.11
Requires-Dist: mcp[cli]<2.0.0,>=1.0.0
Provides-Extra: serve
Requires-Dist: fastmcp>=0.4.0; extra == 'serve'
Requires-Dist: httpx>=0.27.0; extra == 'serve'
Requires-Dist: starlette>=0.41.0; extra == 'serve'
Requires-Dist: uvicorn>=0.30.0; extra == 'serve'
Requires-Dist: x402[evm]>=0.5.0; extra == 'serve'
Description-Content-Type: text/markdown

<!-- mcp-name: io.github.Yarmoluk/ckg-nvidia-nemoclaw -->

# ckg-nvidia-nemoclaw

<!-- Hero image is served from GitHub Pages (docs/og.png), NOT raw.githubusercontent.com.
     This repo is private, so raw.githubusercontent.com 404s for PyPI's anonymous
     fetch and the image renders broken. Pages is public regardless of repo
     visibility. If you move or rename docs/og.png, this URL breaks silently —
     verify with: curl -sI https://yarmoluk.github.io/ckg-nvidia-nemoclaw/og.png -->
<p align="center">
  <a href="https://yarmoluk.github.io/ckg-nvidia-nemoclaw">
    <img src="https://yarmoluk.github.io/ckg-nvidia-nemoclaw/og.png" alt="ckg-nvidia-nemoclaw — NemoClaw as a traversable knowledge graph" width="100%"/>
  </a>
</p>

<p align="center">
  <a href="https://pypi.org/project/ckg-nvidia-nemoclaw/"><img src="https://img.shields.io/pypi/v/ckg-nvidia-nemoclaw?color=0f6e56&label=PyPI" alt="PyPI version"/></a>
  <a href="https://pypi.org/project/ckg-nvidia-nemoclaw/"><img src="https://img.shields.io/pypi/pyversions/ckg-nvidia-nemoclaw?color=0f6e56" alt="Python"/></a>
  <a href="LICENSE"><img src="https://img.shields.io/badge/data-ELv2-0f6e56" alt="Data: ELv2"/></a>
  <a href="LICENSE-CODE"><img src="https://img.shields.io/badge/code-MIT-13201c" alt="Code: MIT"/></a>
  <a href="https://github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf"><img src="https://img.shields.io/badge/F1-0.471_%C2%B73.8%C3%97_RAG-0f6e56" alt="F1: 0.471 · 3.8× RAG"/></a>
  <a href="https://glama.ai/mcp/servers/Yarmoluk/ckg-nvidia-nemoclaw"><img src="https://glama.ai/mcp/servers/Yarmoluk/ckg-nvidia-nemoclaw/badges/score.svg" alt="ckg-nvidia-nemoclaw MCP server"/></a>
</p>

**An auditable knowledge graph for NVIDIA NemoClaw — deterministic agent answers with cryptographic source traceability.**

Every edge traces to a declared relationship and a SHA-256-pinned source document. Built for platform engineers, agent developers, and docs teams who need verifiable answers about NemoClaw dependencies, runtimes, policy, and deployment paths — not model inference.

> Not a general-purpose semantic search layer. If it's not a declared edge, the graph doesn't return it.

```bash
pip install ckg-nvidia-nemoclaw
# or: uvx ckg-nvidia-nemoclaw
```

[PyPI](https://pypi.org/project/ckg-nvidia-nemoclaw/) · [Benchmark paper](https://github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf) · [Interactive graph →](https://yarmoluk.github.io/ckg-nvidia-nemoclaw) · [graphifymd.com](https://graphifymd.com)

---

## What it is

55 nodes · 74 edges · the full NemoClaw stack as a typed dependency graph. Pre-structured, traversable, deterministic. Served over MCP. No inference at query time.

```
get_prerequisites("ManagedMCPServer")

→ ManagedMCPServer
  ├─ [ENABLES]  NemoClaw               ← platform root
  ├─ [REQUIRES] NetworkPolicy          ← root concept, no dependencies
  └─ [REQUIRES] L7Proxy
       ├─ [IMPLEMENTS] OpenShell
       └─ [REQUIRES]   SharedGateway
            └─ [IMPLEMENTS]  InferenceProvider

  269 tokens · declared edges only · no inference
  RAG equivalent: ~2,982 tokens · probabilistic
```

```
query_ckg("ProgressiveToolDisclosure")

← [IMPLEMENTS] OpenClaw
← [IMPLEMENTS] Hermes
← [IMPLEMENTS] LangChain_Deep_Agents

All three runtimes share this mechanism.
RAG returns three separate docs. The graph knows — it's a declared edge.
```

---

## Source provenance — verifiable to the byte

Every node carries a `source_url` and a `source_hash` (SHA-256 of the source document's bytes at extraction time). An edge isn't just *asserted* from a source — it's pinned to a specific version of it.

```bash
# Verify any node's source hasn't changed since extraction
curl -s https://docs.nvidia.com/nemoclaw/latest/ | sha256sum
# expected: 3d5bc97645f1ea274497ee6b931d9649990504daa9fa9ecc56411c324de0beb8
```

**The full audit chain:**

```
edge answer
  → graph commit hash       (git log -- nemoclaw.csv)
  → source_content_hash     (sha256 of page bytes at extraction time)
  → knowledge_source_ref    (URL — fetch hint, not trust anchor)
```

A hash mismatch means either the source changed (stale edge → re-extract) or the graph was patched without re-fetching (silent edit → investigate). No judgment required. Run `scripts/refresh_hashes.py` to recompute.

Via MCP — `verify_source("CorporateCA")`:

```
source_url:  https://docs.nvidia.com/nemoclaw/latest/
source_hash: sha256:3d5bc97645f1ea274497ee6b931d9649990504daa9fa9ecc56411c324de0beb8
verify:      curl -s '<url>' | sha256sum
```

Reference implementation of `knowledge_source_ref` + `source_content_hash` from [GuardrailDecisionV1](https://github.com/crewAIInc/crewAI/issues/4877).

---

## What developers are actually hitting

Signal from 123 GitHub issues, [HN 47427027](https://news.ycombinator.com/item?id=47427027), [HN 47435066](https://news.ycombinator.com/item?id=47435066), and hands-on walkthroughs.

**01 — Context bloat in tool loops.** Agents forget tool schemas after loop iterations. The model re-infers NemoClaw's architecture on every query instead of reading declared structure.

**02 — "Which agent is burning my budget?"** OpenShell makes token spend visible per agent for the first time. The next question is how to reduce it. CKG is that answer.

**03 — The Policy Source Gap.** NVIDIA's own OpenShell knowledge graph names this explicitly: the missing layer between the runtime policy engine and the structured knowledge agents need. We filled it.

---

## Declared relationships, not confidence scores

Every edge was extracted from a source document and given a type. No probabilistic weights, no cosine similarity scores, no confidence intervals. An edge either exists — declared, typed, sourced — or it doesn't. When the answer isn't in the graph, the traversal returns nothing rather than a hallucinated approximation.

**Edge types:**

| Type | Meaning | Example |
|------|---------|---------|
| `REQUIRES` | Hard prerequisite — A cannot function without B | `OpenShell REQUIRES L7Proxy` |
| `ENABLES` | Capability unlock — A makes B possible | `ManagedMCPServer ENABLES NetworkPolicy` |
| `IMPLEMENTS` | Concrete instantiation of an abstract concept | `OpenClaw IMPLEMENTS ProgressiveToolDisclosure` |
| `RELATES_TO` | Conceptual proximity, no dependency direction | `SecurityHardening RELATES_TO Sandbox` |

**Why no confidence levels?** The edge type *is* the confidence signal. `REQUIRES` means load-bearing and sourced; `RELATES_TO` means real but weaker. A missing edge is silence from a source-grounded system — not a soft no, not a low-confidence guess.

```
✗ RAG:  "CorporateCA is probably used for identity management... (similarity: 0.81)"
        Score is on the chunk, not the claim. The claim itself is unverified.

✓ CKG:  "CorporateCA is anchored at image build for TLS interception proxy traversal."
        No score. Declared edge. Traces to security hardening source doc.
```

---

## A/B test — NemoClaw domain, local models, no GPU

30 questions from real GitHub issues · CPU only · Ollama · temperature 0 · seed 42

| Category | Bare model | + CKG | Lift |
|----------|-----------|-------|------|
| Lookup F1 | 0.100 | 0.171 | **+71%** |
| Multi-hop F1 | 0.058 | 0.100 | **+73%** |
| Prereq-chain F1 | 0.077 | 0.156 | **+103%** |
| Key-fact accuracy | 9.3% | 22.3% | **+13pp** |

> phi4-mini and nemotron-mini truncate at ~2,050 tokens. The CKG is 6,837 tokens — only 30% loads. Prereq-chain F1 still doubles on that fraction. Full-context models widen the gap further.

**L01 — lookup:**
```
Q: What are the three agent runtimes in NemoClaw?
✗ Bare: "NemoClaw supports TensorFlow, PyTorch, and ONNX Runtime..." [invented]
✓ CKG:  "OpenClaw (default), Hermes (NEMOCLAW_AGENT=hermes),
         LangChain Deep Agents (NEMOCLAW_AGENT=dcode)" [declared edges, correct]
```

**P08 — prereq-chain (best Δ +0.261):**
```
Q: How does CorporateCA integrate into NemoClaw's security chain?
✗ Bare: "CorporateCA, a cloud-native IAM solution from NVIDIA..." [hallucinated]
✓ CKG:  "CorporateCA is anchored at the image build stage for TLS
         interception proxy traversal." [exact mechanism, correct]
```

**L08 — lookup:**
```
Q: What enterprise manufacturing deployment uses NemoClaw via the FOX Blueprint?
✗ Bare: "FOX (Flexible Open-Source Object Tracking)..." [invented acronym]
✓ CKG:  "Foxconn's MoMClaw is a production deployment of the FOX Blueprint." [correct]
```

---

## Install

**Add to claude.ai (no install required):**

```
https://ckg-nvidia-nemoclaw.onrender.com/mcp
```

Settings → Connectors → Add connector → paste URL.

**Local — Claude Desktop / Claude Code:**

```bash
pip install ckg-nvidia-nemoclaw
# or
uvx ckg-nvidia-nemoclaw
```

```json
{
  "mcpServers": {
    "nemoclaw": {
      "command": "uvx",
      "args": ["ckg-nvidia-nemoclaw"]
    }
  }
}
```

---

## Tools

| Tool | Description |
|------|-------------|
| `ask_nemoclaw(question)` | Natural language query — auto-detects concept, traverses the relevant subgraph |
| `query_ckg(concept, depth)` | Typed subgraph around a specific concept (1–5 hops) |
| `get_prerequisites(concept)` | Full upstream prerequisite chain — every dependency in order |
| `search_concepts(query)` | Fuzzy search across all 55 concepts |
| `list_domains()` | Available domains and node/edge counts |
| `verify_source(concept)` | Source URL + SHA-256 hash for any concept — full audit chain back to source bytes |
| `route_query(question)` | **CKG Router** — returns optimal model tier + reasoning approach derived from graph depth |

### `route_query(question)` — CKG Router

**The graph depth IS the routing signal.** NemoClaw's dependency chains have typed hops that directly signal reasoning complexity. No classifier, no cost heuristic — the graph decides.

| Hop depth | Model | Reasoning |
|---|---|---|
| 1 | haiku | direct |
| 2 | sonnet | generic_cot |
| 3+ | opus | sparql_cot |

**A/B — before vs after routing:**

```
# Baseline (no routing)
query_ckg("L7Proxy", depth=3)
→ subgraph returned, caller guesses model...
→ calls GPT-4o, ~2,400 tokens, $0.072/query

# Treatment (route_query)
route_query("L7Proxy")
→ model_tier: opus
→ reasoning_approach: sparql_cot
→ why: 3-hop chain (OpenShell → CONNECT Proxy → CorporateCA → L7Proxy)
→ structured context injected, ~620 tokens, $0.009/query

Same answer. 79% cheaper. Graph depth made the decision.
```

---

## What's in the graph

**55 nodes · 74 edges · 4 edge types: `REQUIRES` · `ENABLES` · `IMPLEMENTS` · `RELATES_TO`**

| Layer | Concepts |
|-------|----------|
| Agent runtimes | OpenClaw · Hermes (Nous Research) · LangChain Deep Agents |
| Platform | OpenShell · NVIDIA Agent Toolkit · OpenShell TUI · CLI |
| Inference | SharedGateway · vLLM · Ollama · NIM Local · ModelRouter |
| Policy | NetworkPolicy · PolicyTier (Restricted/Balanced/Open) · PolicyPreset · Telegram · Discord · Slack |
| Security | L7Proxy · Landlock LSM · CONNECT Proxy · CorporateCA · SecurityHardening · Sandbox |
| Agent features | Progressive Tool Disclosure · Context Compaction · Heartbeat · Snapshots · Shields |
| Deployment | DGX Spark · DGX Station · macOS Apple Silicon · WSL2 · Brev |
| Ecosystem | FOX Blueprint · MoMClaw (Foxconn) · Nemotron 3 Ultra · Agent Harness |

Every node traces to a source at `docs.nvidia.com/nemoclaw/latest/`, the FOX Blueprint, or the Nemotron 3 Ultra ecosystem docs. Every source is SHA-256 pinned — run `scripts/refresh_hashes.py` to verify.

---

## The dependency graph

> Rendered interactively at [yarmoluk.github.io/ckg-nvidia-nemoclaw](https://yarmoluk.github.io/ckg-nvidia-nemoclaw). On PyPI and plain markdown viewers, the Mermaid source below is human-readable as-is.

```mermaid
graph TD
    NC[NemoClaw] --> OS[OpenShell]
    NC --> OC[OpenClaw]
    NC --> HM[Hermes]
    NC --> LC[LangChain Deep Agents]
    NC --> MCP[ManagedMCPServer]
    NC --> AH[AgentHeartbeat]

    OC --> PTD[ProgressiveToolDisclosure]
    HM --> PTD
    LC --> PTD

    OS --> L7[L7Proxy]
    OS --> LL[LandlockLSM]
    OS --> CP[CONNECT_Proxy]

    L7 --> SG[SharedGateway]
    SG --> IP[InferenceProvider]
    IP --> vLLM[vLLM]
    IP --> OLL[Ollama]
    IP --> NIM[NIM_Local]
    IP --> MR[ModelRouter]

    MCP --> NP[NetworkPolicy]
    MCP --> L7

    NP --> PT[PolicyTier]
    NP --> PP[PolicyPreset]
    PP --> TG[Telegram]
    PP --> DC[Discord]
    PP --> SL[Slack]

    SH[SecurityHardening] --> SB[Sandbox]
    SH --> LL
    SH --> BP[NemoClaw_Blueprint]

    NIM --> DGX[DGX_Spark]

    style NC fill:#0f6e56,color:#fff
    style PTD fill:#1a5c47,color:#fff
    style SH fill:#1a5c47,color:#fff
    style NP fill:#1a5c47,color:#fff
    style IP fill:#2d7a5e,color:#fff
```

---

## Community pulse

123 open GitHub issues · [HN 47427027](https://news.ycombinator.com/item?id=47427027) · [HN 47435066](https://news.ycombinator.com/item?id=47435066)

- **#7084** — Hermes shows ready but tool calls silently fail ("phantom-ready") · async state between OpenShell and agent runtime not synchronized
- **#360** (47↑) — "Can I run local with no API key?" · `inference.local` requires NVIDIA API key even in offline mode
- **#1832** — Multi-sandbox SharedGateway conflicts · two sandbox containers claim the same InferenceProvider slot
- **#2991** — Context window fills after ~12 tool calls in OpenClaw · no auto-compaction
- **#5133** — PolicyPreset Telegram/Discord integration underdocumented

---

## Sources

Every node and edge traces to one of these. No probabilistic inference — declared relationships only.

| Type | Source | Coverage |
|------|--------|----------|
| Official | `docs.nvidia.com/nemoclaw/latest/` | Core platform — agent runtimes, OpenShell, security, inference routing, policy |
| Official | FOX Blueprint docs | MoMClaw (Foxconn) manufacturing deployment, enterprise patterns |
| Official | Nemotron 3 Ultra ecosystem | DGX Spark/Station, ModelRouter, NIM Local integration |
| Official | Security hardening guide | LandlockLSM, L7Proxy, CorporateCA, CONNECT Proxy, Sandbox chain |
| Official | Managed MCP Server docs | ManagedMCPServer, NetworkPolicy, PolicyTier, PolicyPreset, messaging bindings |
| Official | Agent features reference | ProgressiveToolDisclosure, Context Compaction, Heartbeat, Snapshots, Shields |
| Community | GitHub issues (source URL unresolvable — `stale: true`) | 123 issues — phantom-ready, local API key, SharedGateway conflicts. Cited URL 404s; needs re-pointing at the correct upstream tracker before this row is trusted. |
| Community | [HN 47427027](https://news.ycombinator.com/item?id=47427027) · [HN 47435066](https://news.ycombinator.com/item?id=47435066) | Local inference gap (83↑) · token burn visibility |
| Dataset | [huggingface.co/datasets/danyarm/ckg-benchmark](https://huggingface.co/datasets/danyarm/ckg-benchmark) | KRB v0.6.2 — 7,928 queries, 30 NemoClaw-domain questions |
| Benchmark | [github.com/Yarmoluk/ckg-benchmark/paper/main.pdf](https://github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf) | Full methodology, F1 0.471, RAG/GraphRAG baselines |

---

## Benchmark (KRB v0.6.2 locked)

| System | Macro F1 | Mean tokens | Cost / 1k queries |
|--------|----------|-------------|-------------------|
| **CKG** | **0.471** | 269 | $7.81 |
| RAG | 0.123 | 2,982 | $76.23 |
| GraphRAG | 0.120 | ~3,000 | ~$76 |

7,928 queries · 5-hop F1: 0.772 (CKG) vs 0.170 (RAG) · [dataset](https://huggingface.co/datasets/danyarm/ckg-benchmark) · [full paper](https://github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf)

---

## Licensing

Three layers, three licenses, one plain-English answer to each question you actually have:

| Layer | License | Plain English |
|-------|---------|---------------|
| **Server code** — `server.py`, `graph.py`, `serve.py`, `scripts/` | MIT ([LICENSE-CODE](LICENSE-CODE)) | Do anything. Fork it, embed it, sell products built on it. No attribution required. |
| **Graph data** — `domains/nemoclaw.csv` + source hashes | Elastic License 2.0 ([LICENSE](LICENSE)) | Free for all internal and commercial use. The one thing you cannot do: offer this graph as a hosted or managed service that competes with Graphify.md. |
| **Extraction pipeline, benchmark harness, curation methodology** | Proprietary — Graphify.md | Not in this repo. This is the compounding asset — how 97 domains get built and maintained, not the outputs. |

**Can I build an agent or product using this CKG?** Yes. No restrictions, no attribution required.

**Can I run this inside my company's infrastructure?** Yes. ELv2 allows all internal commercial use.

**Can I fork the server code and build my own CKG on a different domain?** Yes. Server code is MIT.

**Can I offer "NemoClaw CKG as a Service" commercially?** No. That's the one thing ELv2 blocks.

**What's actually proprietary?** Not the graph topology — anyone can read the NVIDIA docs. What's proprietary is the extraction process that decides *which* 55 nodes out of thousands matter, *why* L7Proxy `REQUIRES` SharedGateway and not `RELATES_TO`, and the benchmark framework that proves the result is correct.

---

## EVAL

```
benchmark: ckg-benchmark v0.6.2
dataset: huggingface.co/datasets/danyarm/ckg-benchmark
benchmarked: true
this_domain_f1: 0.576
queries_tested: 28
baseline_f1: 0.156
lift_vs_baseline: +269%
model: phi4-mini (Ollama local)
rag_baseline_f1: 0.123
graphrag_baseline_f1: 0.120
mean_tokens: 269
paper: github.com/Yarmoluk/ckg-benchmark/blob/main/paper/main.pdf
```

---

Built by [Graphify.md](https://graphifymd.com) · [97 domains](https://graphifymd.com/pro/) · [PyPI](https://pypi.org/project/ckg-nvidia-nemoclaw/) · patent pending

*Community-built. Not affiliated with, endorsed by, or sponsored by NVIDIA Corporation. NemoClaw is a trademark of NVIDIA Corporation. All referenced trademarks belong to their respective owners.*
