Metadata-Version: 2.4
Name: webpilot-cli
Version: 0.2.8
Summary: pip launcher for the WebPilot browser agent (npm package: @capagents/webpilot)
Author: CLI-Agents
License: MIT
Keywords: webpilot,browser,agent,cli,opentui
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"

# WebPilot

Browser agent CLI. Say what you want in plain language: a goal to carry out in a browser, test cases to run from Jira, Azure DevOps, Xray or a file, a test to write, or a CI pipeline to set up. WebPilot's agent works out what that needs and calls WebPilot's tools. The WebPilot engine drives Chrome, and a run that passes becomes a test in your repo's own language and framework, written by **Scribe**, WebPilot's coding agent.

The engine is the browser service from test-agent-nexus, copied into this package (`py/webpilot_engine`). It keeps the nexus behaviour: error-recovery system prompt, fast-mode agent tuning, Chrome launch flags, the search/select loop breaker, the 600 s run timeout, compact workflow YAML with semantic locators. There is no dependency on test-agent-nexus, FastAPI, or a database.

The shell is [OpenTUI](https://opentui.com). `webpilot run` does the same without the shell, and `webpilot suite` is the fixed command pipelines call. See [HOW_TO_USE.md](HOW_TO_USE.md), the interactive [architecture page](webpilot-architecture.html) (`webpilot docs --open`), and [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md).

## Install

The CLI is published on **npm** as `@capagents/webpilot`. **PyPI** package `webpilot-cli` is a pip launcher for the same `webpilot` command. Either one is the only thing to install, on Windows, macOS or Linux:

```bash
npm install -g @capagents/webpilot
# or
pip install webpilot-cli

webpilot setup   # optional: the first live run does this too
```

Everything else is downloaded on first use into `~/.webpilot/tools` (`%USERPROFILE%\.webpilot\tools` on Windows), checksum verified, without admin rights and without touching PATH or shell profiles:

- **Bun** (the CLI and the OpenTUI shell run on it). A Bun 1.3+ already on PATH or in `~/.bun/bin` is used instead. `WEBPILOT_BUN` points at a specific one.
- **The agent runtime** that WebPilot's agent and Scribe, the coding agent, run on, at the version WebPilot is tested with. `code.runtime` points at a specific binary.
- **uv**, which then fetches **Python** for the engine. A uv already installed is used instead; when uv cannot be downloaded, a Python 3.11+ on PATH (or the Windows `py` launcher) is the fallback.
- **The WebPilot engine** and its Python libraries in `~/.webpilot/engine/.venv`. Installed Google Chrome is used, or Chromium is installed when Chrome is missing. `WEBPILOT_PYTHON` points at an existing Python that already has the engine's libraries; when that Python cannot run the engine (missing, too old, or without them), WebPilot says so and uses its own environment. A broken environment is rebuilt on the next run.

`webpilot setup` does all of this up front; otherwise the first command that needs a piece fetches it. Upgrading is `webpilot update`: it updates every `webpilot` on PATH (bun, npm and pip installs) and waits while the registries catch up with a new release. A release that pins a newer Bun, agent runtime, uv or engine library downloads it on the next run and removes the old one. Deleting `~/.webpilot` removes everything WebPilot downloaded. Behind a proxy, set `HTTPS_PROXY`. The npm launcher's Bun download also needs `NODE_USE_ENV_PROXY=1` (Node 24+); on older Node, use the pip launcher or install Bun yourself.

From this repo:

```bash
cd WebPilot
bun install
bun src/index.ts init
bun src/index.ts start
```

One-shot release (same version on both registries; assumes `npm` and `twine` are already logged in):

```bash
./scripts/publish.sh           # current version
./scripts/publish.sh 0.2.0     # bump + publish
./scripts/publish.sh --dry-run # pack only
```

## Privacy

WebPilot sends no telemetry and turns it off in everything it starts:
- The engine: PostHog telemetry, cloud sync, and the price-list download are off, and the blank-tab logo its library loads from a CDN is replaced by WebPilot's start page, which loads nothing.
- The agent runtime: auto-update, session sharing, OpenTelemetry spans, and the model catalogue download are off; it uses its bundled model list.
- The repo's test command: `DO_NOT_TRACK=1`, plus the opt-outs for Next.js, Nuxt, Astro, Gatsby, Storybook, Turborepo, Angular, .NET, and Cypress.

`OTEL_EXPORTER_OTLP_*` variables are removed from those processes, and your environment cannot turn any of this back on. The only traffic is to your model endpoint and the sites the agent visits. The engine also downloads its ad-block and cookie-banner extensions from the Chrome Web Store once.

## Configure

`init` writes:

| File | Purpose |
|------|---------|
| `webpilot.yaml` | Browser, prompts, export, active profile |
| `llms.json` | Named LLM profiles (azure, openai, ollama, openai_compatible, mock) |
| `.env.example` | API key names |

```bash
webpilot models
webpilot run "Read the homepage" --url https://example.com --profile mock --plain
```

`mock` is an offline demo on a synthetic page: no engine, browser, or model. Pick a live profile (Tab in the shell, or `llm.active`) for a real run.

## CLI

```bash
webpilot start
webpilot start -g "Go to booking.com and search hotels in Mumbai" -p azure-gpt4o

webpilot run "Get a quote on example.com for a family of four" --profile azure-gpt4o --headed
webpilot run "Run SHOP-12 and SHOP-14, 2 browsers at once"
webpilot run "Write a Playwright test for the login flow on staging.app.test" --repo ../app
webpilot run "Test the refund rules on https://acme.atlassian.net/wiki/spaces/QA/pages/123"
webpilot login shop                      # sign in by hand once; runs reuse the session
webpilot run --direct "Read the homepage" --url https://example.com --out ./.webpilot/reports/demo   # one browser goal, no agent
```

There are no commands to pick a source or a mode. WebPilot's agent reads the request and decides: a browser goal, test cases (Jira keys, Azure DevOps ids or test plans, Xray, TestRail, Zephyr Scale or qTest cases, files in the working directory, or cases pasted into the message), requirements to turn into cases (a Confluence page, a story), a test, or a pipeline. When a reference could mean more than one thing, it asks. The agent picks the site to open; `--url` only pins a start page.

Inside the shell, type the request and press Enter. The only commands are for the session: `/stop`, `/headed`, `/profile`, `/clear`, `/help`, `/exit`. Everything the other commands do can be asked for in the shell too, since each one is a tool the agent calls: "heal the checkout tests", "check the smoke cases once", "monitor them every 10 minutes", "sign in to shop", "forget the shop session", "which profiles do I have?". Results appear as cards in the transcript, and when a step needs you (signing in by hand), a **Your turn** prompt waits until you pick **Done, continue** or **Cancel**.

WebPilot's agent needs the agent runtime (downloaded automatically, see Install) and a live profile. It uses the active `llms.json` profile, or `code.model: provider/model`. With the `mock` profile, or without the agent runtime, the request runs as one browser goal.

`webpilot run` exits 0 when the job is done, 1 when it failed, 2 when it was stopped, and 3 when the agent needs more information.

## Page memory

Every live run records each page it visits in `.webpilot/memory` next to `webpilot.yaml` (`~/.webpilot/memory` when there is no config file): every interactive element, with all of its locators (test id, id, role and name, label, placeholder, alt, title, href, text, CSS, XPath), a confidence % for each, and when it was first seen, last seen, and last used. Confidence rises when a locator keeps matching exactly one element and works when used, and falls when it goes missing or fails.

A goal that passed before is replayed from memory with no model calls. Each step finds its element by the highest-confidence locator that still matches, so a renamed button or a changed id heals itself. The agent then checks the result with one model call (`replay: verify`), or the run finishes without the model (`replay: trust`). If a step cannot be found, the agent takes over from that point. Generated specs use the best locator from memory.

```bash
webpilot memory                          # hosts, pages, elements, flows
webpilot memory show booking.com         # a host's pages
webpilot memory show https://app.test/login --locators   # every locator with confidence
webpilot memory flows --steps            # recorded goals and their steps
webpilot memory clear booking.com --yes
webpilot run "..." --replay trust       # or --replay off, --no-memory
```

## Tests

Point WebPilot at a repo and a run that passes becomes a test in that repo. The agent gets the run's steps and the locators and confidence from page memory. It studies the repo's framework, page objects, fixtures, and helpers, reuses them, and writes the test in the repo's language and test framework (C# NUnit, TypeScript, Python, Java, …). The browser agent records what it checked with the exact text on the page, and the test asserts those checks with locators the Playwright replay found. WebPilot runs the test command and hands any failure back (the error first, the page dump saved to a file), until the test passes or the attempts run out. What Scribe learned about the repo is kept, so the next job there starts from its notes instead of reading the repo again. The test counts only when WebPilot's own run of it passes.

```bash
webpilot run "Log in and open the invoices page on staging.app.test" --repo ../app
webpilot code --repo ../app          # write a test for the latest run
webpilot code .webpilot/reports/20260928-172700 --repo ../app --attempts 3
```

In the shell and `webpilot run`, every passed run becomes a test unless you say otherwise: in `--repo`, in `code.repo` from `webpilot.yaml`, or else in the current folder. `--no-code` turns tests off. The transcript shows the handover to Scribe, the coding agent, the files it reads and writes, and WebPilot's run of the test, and each request ends with a card listing the run's report and the test with its result (or why no test was written).

### Fixing broken tests

When the app changes, tests break. `webpilot heal` runs the repo's tests, and for each failing test the agent re-runs its flow in the real browser. If the flow still works, the test is out of date: the agent updates its steps and locators (in the page object, when that's where they live) from what the browser just did and from page memory. If the browser fails at the same step, the app is broken there: the test is left alone and the failure is reported as an app bug. Tests are never deleted, skipped, loosened, or given longer timeouts to make them pass.

WebPilot re-runs every healed test and then the whole command itself; a test counts as healed only when WebPilot's own run of it passes.

```bash
webpilot heal --repo ../app                              # code.test_command, or the agent finds the command
webpilot heal "npx playwright test tests/checkout" --repo ../app
webpilot heal -- npx playwright test --project=chromium
webpilot run "The login tests broke after the redesign, fix them" --repo ../app   # the agent heals them the same way
```

The report is `heal.md` (and `heal.json`) in `.webpilot/reports/heal/<time>/`, with each test's verdict, the files changed, the browser runs, and the test output before and after. `webpilot heal` exits 0 when the tests pass (or nothing needed healing), 1 when some still fail, and 2 when stopped.

### Running existing tests

"Run all tests related to login" or "run the tests added this sprint" finds the repo's existing tests and runs them once you agree. Scribe reads the tests (and git history for a date range or sprint) and decides each one from what it does, not from its name. WebPilot runs the framework's list command to confirm the selection and shows the list. Nothing runs until you pick **Run**, and then WebPilot runs exactly those tests and shows a pass or fail for each. For a sprint, WebPilot fetches the sprints from Azure DevOps and Jira and asks which one you mean.

```bash
webpilot run "Run all test cases related to login" --repo ../app
webpilot run "Which tests were added this sprint?" --repo ../app
```

## Test cases and suites

Run existing test cases instead of typing goals: ask for them by name ("run SHOP-12", "run test plan 12 suite 34", "run the cases in smoke.yaml") or paste them into the message. Each case becomes one goal: its steps are what the browser agent does and its expected results are what it checks, and the agent's own verdict decides pass or fail. With a repo set, every passed case also gets a test.

`webpilot suite` runs cases the same way every time, with no agent choosing anything, which is what pipelines need. It takes explicit references:

```bash
webpilot cases cases.yaml                     # what a source gives, without running anything
webpilot suite cases.yaml --parallel 3        # run them, 3 browsers at once
webpilot suite ado:plan=12/suite=34 --repo ../app
webpilot suite jira:jql="project = SHOP AND labels = smoke"
webpilot suite xray:plan=SHOP-100
webpilot suite "Open example.com and check the title says Example Domain"
```

| Source | Reference |
|--------|-----------|
| Plain text | the text itself, quoted |
| Text or Markdown file | `.txt` / `.md`, one case per block between `---` lines |
| CSV | `.csv` with `id, title, url, description, preconditions, tags, action, data, expected` columns |
| YAML / JSON | `.yaml` / `.json`, a list of `{id, title, url, preconditions, steps: [{action, data, expected}]}` |
| Gherkin | `.feature`, one case per scenario (and per Examples row) |
| Azure DevOps | `ado:123,456` (test case ids), `ado:plan=12`, `ado:plan=12/suite=34` |
| Jira | `jira:SHOP-1,SHOP-2`, `jira:jql=<query>` (Cloud and Server/Data Center) |
| Xray | `xray:SHOP-1`, `xray:jql=<query>`, `xray:plan=SHOP-100` (Cloud and Server/Data Center) |
| TestRail | `testrail:C12,C13`, `testrail:run=45`, `testrail:plan=7`, `testrail:suite=3` |
| Zephyr Scale (Cloud) | `zephyr:SHOP-T1,SHOP-T2`, `zephyr:cycle=SHOP-R5`, `zephyr:folder=12` |
| qTest | `qtest:1234` (test case ids), `qtest:cycle=88`, `qtest:suite=9` |

For a case with numbered steps the browser agent reports a verdict for every step (passed, failed, or not run, with what it actually saw), not only for the case.

Results go back where the cases came from, with the step verdicts: an Azure DevOps test run for a plan (or a comment on the work item), a Jira comment, an Xray Test Execution with the failure screenshot, a TestRail run (the run the cases came from, or a new one) with step results and the screenshot, a Zephyr Scale test cycle, or qTest test logs. `--no-write-back` or `sources.write_back: false` turns it off.

### Bugs

With `sources.bugs`, each failed case gets a bug in Jira or Azure DevOps with the steps and verdicts, where it ended, the build link, and the last screenshot, linked to the test case. WebPilot tags the bug with a fingerprint of the case, so the next failure adds a "still failing" comment to the open bug instead of filing another, and a pass adds a note that it passes again.

```yaml
sources:
  bugs: { system: jira, project: SHOP }          # or system: azure_devops (uses azure_devops.project)
```

### Reports

Every suite writes `report.html`, `junit.xml`, and `suite.json` to `.webpilot/reports/suites/<time>/`, with each case's run folder beside them; every single run writes its own `report.html` too. The report is one HTML file with the screenshots embedded, so you can open it straight from disk, attach it to a build, or mail it; no server is needed. It has an overview (verdict, totals, step checks, tokens, cost, page memory savings, browser errors, accessibility issues, bugs filed, what changed since the last run, results over earlier runs, the cases that need attention, results by source and tag, the slowest cases, and what was posted where), a test results page (each case's step table with expected and actual results, why it failed, its last passing screenshot beside this run's, browser errors per step, page memory with healed locators, the browser agent's steps with a screenshot each and the clicked element outlined, a timing waterfall, accessibility issues, and the trace and workflow to read in place), a failures page, since last run, trends and flaky tests from the earlier suite runs in the same `.webpilot/reports/suites/` folder, time and cost (model vs browser time, a parallel worker timeline, cost per case from the model's `price` in `llms.json`), coverage by requirement and group from the test system, accessibility checks for every page visited, and the environment (model, browser, WebPilot version, CI build, branch, commit). It exports Markdown and JSON and prints cleanly. See `samples/reports/nightly/20260929-020005/report.html` for an example. `report.screenshots: failures` keeps screenshots for failed runs only, `off` drops them; `report.accessibility: false` skips the page checks. On GitHub Actions the suite also writes a job summary.

## Signing in

Sites that need a login go under `logins:`. The browser agent types credentials from the environment without ever seeing them (it only sees placeholders such as `shop_password`), fills authenticator codes from a TOTP secret, and keeps the signed-in session in `.webpilot/sessions/` so the next run starts signed in. Replays read the same environment variables, and so do the tests the coding agent writes.

```yaml
logins:
  shop:
    url: https://staging.shop.test/login
    username_env: SHOP_USER
    password_env: SHOP_PASSWORD
    totp_env: SHOP_TOTP_SECRET      # optional
```

Credentials in a message or a test case work too: WebPilot's agent remembers them for the session as a login for the site and its sign-in hosts, so goals and reports never contain the password, the exact value is typed, and follow-up runs start signed in.

For single sign-on, captchas, or security keys, sign in by hand once: `webpilot login shop` opens a browser, you sign in, press Enter, and the session is saved. `webpilot logins` lists logins and sessions; `webpilot logins clear shop` forgets one. Session files are private to your user and have their own `.gitignore`.

## Azure DevOps, Jira, and other systems

WebPilot's agent connects to each system's MCP server, configured from the same `sources:` settings in `webpilot.yaml`: organization, project, site URL, and the names of the environment variables that hold the tokens. You don't write any MCP config yourself.

| System | MCP server | Needs |
|--------|------------|-------|
| Azure DevOps Services | Microsoft's [Azure DevOps MCP server](https://github.com/microsoft/azure-devops-mcp) (`npx @azure-devops/mcp`), with work items, boards, test plans, pipelines, and repos | Node.js, and `ADO_PAT` or `az login` |
| Jira Cloud, Server/Data Center | [mcp-atlassian](https://github.com/sooperset/mcp-atlassian) (`uvx`), with issues, JQL search, comments, and transitions. `mcp: rovo` uses [Atlassian's hosted server](https://github.com/atlassian/atlassian-mcp-server) instead (Cloud, and an admin must allow API tokens) | uv; `JIRA_EMAIL` and `JIRA_API_TOKEN` (Cloud) or a personal access token (Server/DC) |
| Confluence | mcp-atlassian, shared with Jira (or its own when Jira isn't set up); Atlassian's hosted server covers it on the same site | `confluence: true` on a Jira Cloud site, or `confluence.url` + tokens |
| Xray, TestRail, Zephyr Scale, qTest | none; WebPilot's API | `XRAY_CLIENT_ID` / `XRAY_CLIENT_SECRET`, `TESTRAIL_EMAIL` / `TESTRAIL_API_KEY`, `ZEPHYR_API_TOKEN`, `QTEST_TOKEN` |
| Anything else | servers you add under `mcp:` | |

```yaml
sources:
  azure_devops: { org_url: https://dev.azure.com/acme, project: Shop }   # token in ADO_PAT
  jira: { url: https://acme.atlassian.net }                              # JIRA_EMAIL + JIRA_API_TOKEN
mcp:
  github:
    url: https://api.githubcopilot.com/mcp/
    headers: { Authorization: "Bearer {env:GITHUB_TOKEN}" }
```

Tokens never go into the config: servers get them through `{env:NAME}` references that the agent runtime fills in, and the agent never sees them. `webpilot mcp list` shows what your settings give. When the agent starts, the shell and `webpilot run` print which servers connected.

Whatever has no MCP server, or whose server didn't connect, goes through WebPilot's own tools, which call the REST APIs directly: `find_test_cases` and `run_test_cases` for test cases and posting results back, `read_confluence` for Confluence pages, and `query_api` to read anything else from Azure DevOps, Jira, Xray, or Confluence. `query_api` only reads (GET requests and Xray GraphQL queries) and only reaches the configured hosts. `sources.azure_devops.mcp: false` or `sources.jira.mcp: false` uses the API only.

Requirements work too: "write and run test cases for the checkout rules page in Confluence" makes the agent read the page, write one case per rule or acceptance criterion (with the main negative paths), show them, and run them. It saves them into a test system only when you ask.

The agent creates or changes items (a bug, a comment, a status change) only when you ask it to; bugs for failed cases come from `sources.bugs`, not from the agent.

## Pipelines

Ask for the pipeline you want and the agent builds it:

```bash
webpilot run "Add a GitHub Actions workflow that runs the smoke cases in cases.yaml on every pull request"
webpilot run "Set up an Azure DevOps pipeline that runs test plan 12 every night and posts results back"
webpilot run "The WebPilot pipeline on main is failing, fix it"
```

It works in the repo (the working directory, or `--repo`): it reads the repo's existing CI, writes or updates the workflow so it installs Bun, uv and WebPilot, calls `webpilot suite <source...> --no-code --plain`, and publishes `junit.xml` and the report, and validates it (actionlint when it is installed). Then it commits to a new branch named `webpilot/ci-<topic>`, stores the keys the pipeline needs as secrets, triggers the run with `gh` or `az`, watches it, and fixes failures until the run passes (up to `code.max_attempts` failed runs).

Guard rails:
- Pushing goes through WebPilot's `push_branch` tool, which refuses the default branch, `main`, and `master`, and never forces. `git push` itself is blocked for the agent.
- Secrets go through `set_ci_secret`: the agent names an environment variable and WebPilot passes its value straight to `gh secret set` or `az pipelines variable`. The agent never sees the value.
- It needs `gh` (GitHub, logged in) or `az` with the `azure-devops` extension (Azure DevOps, logged in). Without them it still writes, validates, and commits the pipeline, then tells you what to run.

## Monitoring

`webpilot monitor` runs checks on a schedule and tells your team when one breaks. A check is any test case `webpilot suite` takes: a YAML file, a Jira key, an Azure DevOps plan. Each round replays the flows from page memory, so a flow that passed before needs one model call to confirm the result (or none with `replay: trust`).

```yaml
monitor:
  checks: [checks/smoke.yaml, "jira:jql=labels = monitor"]
  every: 15m
  fail_after: 2           # two failing rounds in a row before the first alert, to ride out a flaky run
  remind_every: 2h        # repeat while it keeps failing
  alerts:
    teams: { webhook_env: TEAMS_WEBHOOK_URL }
    email: { host: smtp.office365.com, port: 587, to: [qa-team@example.com] }   # SMTP_USERNAME, SMTP_PASSWORD
```

```bash
webpilot monitor --test-alerts      # a sample alert, to check the webhook and SMTP settings
webpilot monitor                    # every 15 minutes until Ctrl+C
webpilot monitor --once             # one round, for cron or a scheduled pipeline
```

An alert goes out when a check starts failing, again every `remind_every` while it fails, and once when it recovers. One round sends one message covering every check that changed. Teams gets an Adaptive Card (both incoming webhooks and Workflows webhooks accept it) and email gets the failing checks' screenshots attached. Each alert links to the round's report: the local file, or `<report_url>/<round>/report.html` when you publish `.webpilot/reports/monitor/` somewhere. A round that cannot run at all (a test system is down, the engine will not start) alerts as the check "WebPilot monitor".

Rounds go to `.webpilot/reports/monitor/<time>/`, the same as suites, so the report shows trends and flaky checks over the rounds; the last `keep` rounds (100) are kept. `state.json` remembers which checks are failing and since when, so with `--once` in a pipeline, cache that folder between runs. Only one monitor can use a folder at a time. `--once` exits 0 when every check passed and 1 otherwise. Results are posted back to the test system only with `write_back: true`. In the shell, "monitor the smoke cases every 10 minutes" runs the rounds in the background while you keep working, and each round shows as a card; they stop when the shell closes.

## MCP

`webpilot mcp` serves WebPilot's tools over stdio for other agents such as Claude Code, Cursor, or Codex:

```json
{ "mcpServers": { "webpilot": { "command": "webpilot", "args": ["mcp", "--repo", "/path/to/app"] } } }
```

Tools: `browser_run`, `find_test_cases`, `run_test_cases`, `run_details`, `run_test`, `page_memory`, `read_confluence`, `query_api`, `push_branch`, `set_ci_secret`, `find_tests`, `run_tests` (runs only with `confirmed: true`, once the user agreed), `list_sprints`, and one tool per command: `write_test`, `heal_tests`, `monitor`, `logins`, `webpilot_setup`. Browser runs and suites log their progress to stderr, and long calls send progress notifications.
