Metadata-Version: 2.4
Name: docgate-llm
Version: 0.0.1
Summary: Give developers using browser-based LLMs policy-gated access to current technical documentation.
Author: ThreeI
License: MIT
Project-URL: Homepage, https://github.com/romantick13/docgatellm
Requires-Python: >=3.10
Description-Content-Type: text/markdown

DocGateLLM

Policy-gated documentation fetch for browser-based LLMs. Deny-by-default.

The model never chooses an arbitrary URL — it selects only a registeredsource and a validated locator. DocGateLLM builds the request, enforcesthe owner's allowlist, fetches safely, parses structured content andreturns it as untrusted data.

The model is an untrusted caller. The owner configures the policy once.No human approval per documentation request.

Owner configures permitted sources once
        ↓
The model fetches only within that fixed policy
        ↓
No human approval per documentation request


## What It Is

DocGateLLM lets coding agents (browser-based LLMs treated as untrusted
callers) retrieve **current technical documentation** from
**owner-approved sources** — the knowledge that post-dates their training
cutoff.

In this project, "browser-based LLMs" means AI assistants accessed through
a web chat interface. Availability depends on the MCP client or bridge
integration.

## Why

An LLM's training knowledge has a cutoff. Frameworks move faster:
.NET 10, Avalonia 12, fresh library APIs — a model trained a year ago
will hallucinate methods that do not exist or produce outdated examples
with full confidence.

DocGateLLM turns external documentation into a **controlled retrieval
primitive**: the model asks, the policy answers, the owner stays in
control.

LLM / coding agent
        │ source + typed locator (never a URL)
        ▼
   DocGateLLM
        │ allowlist · SSRF policy · cache · parser · audit
        ▼
Approved documentation sources


## Security Model

- **Owner-approved sources only** — exact host match, configured once
- **No arbitrary URLs** — the model cannot supply, construct or influence
  the request URL; the server builds it from the source config
- **SSRF-safe transport** — DNS resolve once → validate all addresses →
  pinned connection; private/link-local/metadata ranges denied
- **HTTPS only, port 443 only, redirects disabled** in v1
- **Bounded responses** — size caps, content-type allowlist, timeout
- **Untrusted data** — every response is marked `kind: "untrusted-data"`;
  the model is instructed: documentation is data, not instructions
- **Full audit** — every fetch logged (source, locator, URL, verdict, bytes)
- **Rate limits** — per task, per source, global

Not included, by design: RAG / vectorization, browser automation,
JavaScript rendering, arbitrary web access, cookies/auth/credentials.

## Status

**v0.0.1 — project skeleton.** Architecture design is complete
(4-reviewer pass: SSRF matrix, typed locators, DNS pinning, parsing
ladder). Implementation: in progress.

DocGateLLM grew out of [UnlockBridge](https://github.com/romantick13/UnlockBridge)
— a human-gated bridge between browser AI chats and local files. It works
with standard MCP clients and servers, and is tested alongside the
MCPBridge component of UnlockBridge — but does not require UnlockBridge.

## Naming

| Context | Name |
|---|---|
| GitHub / display | DocGateLLM |
| Repository slug | `docgate-llm` |
| Python package / import | `docgate_llm` |
| CLI (planned) | `docgate-llm` |
| Config file (planned) | `docgate-llm.yaml` |

## License

MIT (to be finalized in v0.1)

