Ghost Developer Studio

Ghost Hands

Playwright drives. Stagehand thinks. Browser Use wanders.
Ghost Hands answers for every move.

Agent hands for the web — with a governor attached. Every action is classified before it runs, judged against a policy, written to a provenance trail before it executes, and gated on a human approval when it matters.

$ pip install ghost-handsNow on GitHub

v0.3.0 is live at github.com/littlestjames82-sys/ghost-hands; the PyPI listing is being registered and lands at launch. Nothing here is vaporware: every number below was measured on the current build.

The wedge

Every browser-agent tool in the field will click Pay, Send, or Delete the moment a model tells it to. Ghost Hands is the layer that says: classify first.

🛡️

Governed by default

Each action is classified and checked against a policy before the driver moves. Consequential actions stop for approval — and with no approver attached, "ask" means denied. Silence is never consent.

readonly write consequential
🧾

Recorded, then replayable

A JSONL provenance trail captures perceive → decide → govern → execute → result for every step. Any finished run graduates into a deterministic, model-free script: explore with a model once, re-run forever for free.

👻

Our own stack

The browser body is a from-scratch, standard-library CDP client driving Chromium directly. Zero runtime dependencies — no Playwright, no Selenium, nobody else's automation layer under the hood. (We drive Chromium, the browser; we didn't write a browser engine.)

👀

Cheap eyes

Perception is a numbered element map parsed from the page — no screenshot firehose, no vision model. The bundled demo reads a whole page in about 174 tokens (696 map chars).

🩹

Self-healing targets

Pages mutate between looking and clicking. When a target's number no longer points at the element the decider meant, the hands re-find it by descriptor, retry once, and write the heal into the trail.

🧠

Any brain, or none

Pluggable deciders: explicit scripts, deterministic offline rules, or any OpenAI-compatible model (key from the environment only, never stored). Plus tabs, session save/load, real screenshots, and a stdlib MCP server.

Verified, not vibes

Measured on the v0.3.0 build, October 8, 2026 — the same suites that ship in the repo.

170
pytest tests passing
82/82
bench cases (offline suite)
2/2
live cases: example.com + a real Wikipedia search
0
runtime dependencies
perceive  step 4 · 21 elements · 1,842 map chars
decide    step 4 · {"kind": "click", "target": 7}
govern    step 4 · classification=consequential · outcome=ask
execute   step 4 · [7] <button> "Place order — pay $42"  ← held for approval
stop      denied · approval required, no approver
The honest limits: the LLM decider's wire protocol is proven against a local stub endpoint; live-model driving quality is unproven and depends on the model you bring. The Chromium driver is proven on fixture pages plus example.com and Wikipedia — not on the whole web. There are no head-to-head speed benchmarks against other tools, and we won't claim any we haven't run.

Where it's headed

One pair of hands, more bodies: the same action protocol gains an Android body next (via MrGhosty, our phone agent), and a GhostBus transport so other agents can task the hands. Policy packs from Agent Seatbelt / GhostGuard plug into the same governor.