Ghost Developer Studio

Ghost Hands

Playwright drives. Stagehand thinks. Browser Use wanders.
Ghost Hands answers for every move.

Agent hands for the web — with a governor attached. Every action is classified before it runs, judged against a policy, written to a provenance trail before it executes, and gated on a human approval when it matters.

$ pip install ghost-handsLive on PyPI

v0.6.0 is live on PyPI and GitHub; the current build is v0.9.0 — the field kit wave: `ghost-hands doctor` (whole-stack diagnostics with plain-language fixes) and `ghost-hands phone-proof` (the guided on-device proof), alongside trail → HTML audit reports (`ghost-hands report`), MCP parity across all four bodies with phone approvals, the seatbelt policy pack denying prompt-injection-shaped targets, the model-evaluation harness, Android gestures in MrGhosty v1.8.0, approvals by phone, GhostBus agent mode, the Body Protocol, and policy packs. Nothing here is vaporware: every number below was measured on the current build.

The wedge

Every browser-agent tool in the field will click Pay, Send, or Delete the moment a model tells it to. Ghost Hands is the layer that says: classify first.

🛡️

Governed by default

Each action is classified and checked against a policy before the driver moves. Consequential actions stop for approval — and with no approver attached, "ask" means denied. Silence is never consent.

readonly write consequential
🧾

Recorded, then replayable

A JSONL provenance trail captures perceive → decide → govern → execute → result for every step. Any finished run graduates into a deterministic, model-free script: explore with a model once, re-run forever for free.

👻

Our own stack

The browser body is a from-scratch, standard-library CDP client driving Chromium directly. Zero runtime dependencies — no Playwright, no Selenium, nobody else's automation layer under the hood. (We drive Chromium, the browser; we didn't write a browser engine.)

👀

Cheap eyes

Perception is a numbered element map parsed from the page — no screenshot firehose, no vision model. The bundled demo reads a whole page in about 174 tokens (696 map chars).

🩹

Self-healing targets

Pages mutate between looking and clicking. When a target's number no longer points at the element the decider meant, the hands re-find it by descriptor, retry once, and write the heal into the trail.

🧠

Any brain, or none

Pluggable deciders: explicit scripts, deterministic offline rules, or any OpenAI-compatible model (key from the environment only, never stored). Plus tabs, session save/load, real screenshots, and a stdlib MCP server.

🚌

Works the bus (new in 0.4)

ghost-hands bus-agent joins a GhostBus workspace as the agent other agents task: it claims governed work, posts progress, uploads its trail as a shared file — and when the governor asks, the approval request itself travels over the bus. Policy packs (readonly / standard / strict) tune the leash per deployment.

readonly standard strict
📱

One hands, many bodies

A written Body Protocol defines the driver contract, proven by three bodies — real Chromium, an in-memory web, and a simulated phone with apps, notes, toggles, and rendered PNG screenshots — all driven by the same runner, governor, and trail. The Android driver is specified in the protocol; the simulator is its rehearsal.

Verified, not vibes

Measured on the v0.9.0 build, October 8, 2026 — the same suites that ship in the repo.

395
pytest tests passing
95/95
bench cases (offline suite)
2/2
live cases: example.com + a real Wikipedia search
0
runtime dependencies
perceive  step 4 · 21 elements · 1,842 map chars
decide    step 4 · {"kind": "click", "target": 7}
govern    step 4 · classification=consequential · outcome=ask
execute   step 4 · [7] <button> "Place order — pay $42"  ← held for approval
stop      denied · approval required, no approver
The honest limits: the LLM decider's wire protocol is proven against a local stub endpoint, and v0.7.0's eval harness now grades any model you bring on 8 local tasks (the stub's 8/8 proves the harness, not a model); live-model driving quality on real websites is unproven and depends on the model you bring. The Chromium driver is proven on fixture pages plus example.com and Wikipedia — not on the whole web. There are no head-to-head speed benchmarks against other tools, and we won't claim any we haven't run. The bus agent is proven against the real GhostBus server booted locally (both its shapes); the Android driver and the MrGhosty bridge are built and proven against the bridge wire contract — proof on a physical phone is still ahead.

Where it's headed

v0.4.0 shipped the GhostBus transport (other agents can task the hands, and approvals route back over the bus), policy packs (readonly / standard / strict), and the Body Protocol with a simulated phone body. v0.5.0 built the Android body itself — MrGhosty's accessibility service hosts the bridge and AndroidDriver speaks the same protocol, so one pair of hands works the phone and the web. v0.6.0 added the phone approver: when a run hits a consequential action, the question goes to Ryan's phone and only his tap answers it — the bridge API can ask, but it can never approve. v0.7.0 built the measuring instrument for model driving (the eval harness), taught the Android body gestures, and closed the GhostBus file-store loop. Next: proving it on Ryan's phone.