Metadata-Version: 2.4
Name: webpilot-cli
Version: 0.1.9
Summary: pip launcher for the WebPilot browser agent (npm package: @capagents/webpilot)
Author: CLI-Agents
License: MIT
Keywords: webpilot,browser,agent,cli,opentui
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"

# WebPilot

Browser agent CLI. Say what you want in plain language: a goal to carry out in a browser, test cases to run from Jira, Azure DevOps, Xray or a file, a test to write, or a CI pipeline to set up. WebPilot's agent ([opencode](https://opencode.ai)) works out what that needs and calls WebPilot's tools. [browser-use](https://github.com/browser-use/browser-use) 0.13.10 drives Chrome, and every run leaves a replayable Playwright spec.

The engine is the browser-use service from test-agent-nexus, copied into this package (`py/webpilot_engine`). It keeps the nexus behaviour: error-recovery system prompt, fast-mode agent tuning, Chrome launch flags, the search/select loop breaker, the 600 s run timeout, `browser-use-v2-compact` workflow YAML with semantic locators, and the Playwright codegen. There is no dependency on test-agent-nexus, FastAPI, or a database.

The shell is [OpenTUI](https://opentui.com). `webpilot run` does the same without the shell, and `webpilot suite` is the fixed command pipelines call. See [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) and [HOW_TO_USE.md](HOW_TO_USE.md).

## Install

OpenTUI needs **Bun 1.3+**. Node cannot load the native renderer. Live runs also need **uv** or **Python 3.11+** for the browser-use engine.

The CLI is published on **npm** as `@capagents/webpilot`. **PyPI** package `webpilot-cli` is a pip launcher for the same `webpilot` command.

```bash
curl -fsSL https://bun.sh/install | bash
curl -LsSf https://astral.sh/uv/install.sh | sh   # recommended for the engine

npm install -g @capagents/webpilot
# or
pip install webpilot-cli

webpilot setup   # optional: the first live run does this too
```

`webpilot setup` creates `~/.webpilot/engine/.venv` with browser-use 0.13.10. It uses installed Google Chrome, or installs Chromium when Chrome is missing. `WEBPILOT_PYTHON` points WebPilot at an existing Python that already has browser-use.

From this repo:

```bash
cd WebPilot
bun install
bun src/index.ts init
bun src/index.ts start
```

One-shot release (same version on both registries; assumes `npm` and `twine` are already logged in):

```bash
./scripts/publish.sh           # current version
./scripts/publish.sh 0.2.0     # bump + publish
./scripts/publish.sh --dry-run # pack only
```

## Privacy

WebPilot sends no telemetry and turns it off in everything it starts:
- browser-use: PostHog telemetry, cloud sync, and the price-list download are off, and its blank-tab logo (from cf.browser-use.com) is replaced by WebPilot's start page, which loads nothing.
- opencode: auto-update, session sharing, OpenTelemetry spans, and the models.dev download are off; opencode uses its bundled model list.
- The repo's test command: `DO_NOT_TRACK=1`, plus the opt-outs for Next.js, Nuxt, Astro, Gatsby, Storybook, Turborepo, Angular, .NET, and Cypress.

`OTEL_EXPORTER_OTLP_*` variables are removed from those processes, and your environment cannot turn any of this back on. The only traffic is to your model endpoint and the sites the agent visits. browser-use also downloads its ad-block and cookie-banner extensions from the Chrome Web Store once.

## Configure

`init` writes:

| File | Purpose |
|------|---------|
| `webpilot.yaml` | Browser, prompts, export, active profile |
| `llms.json` | Named LLM profiles (azure, openai, ollama, openai_compatible, mock) |
| `.env.example` | API key names |

```bash
webpilot models
webpilot run "Read the homepage" --url https://example.com --profile mock --plain
```

`mock` is an offline demo on a synthetic page: no engine, browser, or model. Pick a live profile (Tab in the shell, or `llm.active`) for a real run.

## CLI

```bash
webpilot start
webpilot start -g "Go to booking.com and search hotels in Mumbai" -p azure-gpt4o

webpilot run "Get a quote on example.com for a family of four" --profile azure-gpt4o --headed
webpilot run "Run SHOP-12 and SHOP-14, 2 browsers at once"
webpilot run "Write a Playwright test for the login flow on staging.app.test" --repo ../app
webpilot run "Test the refund rules on https://acme.atlassian.net/wiki/spaces/QA/pages/123"
webpilot login shop                      # sign in by hand once; runs reuse the session
webpilot run --direct "Read the homepage" --url https://example.com --out ./out/demo   # one browser goal, no agent
```

There are no commands to pick a source or a mode. WebPilot's agent reads the request and decides: a browser goal, test cases (Jira keys, Azure DevOps ids or test plans, Xray, TestRail, Zephyr Scale or qTest cases, files in the working directory, or cases pasted into the message), requirements to turn into cases (a Confluence page, a story), a test, or a pipeline. When a reference could mean more than one thing, it asks. The agent picks the site to open; `--url` only pins a start page.

Inside the shell, type the request and press Enter. The only commands are for the session: `/stop`, `/headed`, `/profile`, `/clear`, `/help`, `/exit`. Everything the other commands do can be asked for in the shell too, since each one is a tool the agent calls: "heal the checkout tests", "check the smoke cases once", "monitor them every 10 minutes", "sign in to shop", "forget the shop session", "which profiles do I have?". Results appear as cards in the transcript, and when a step needs you (signing in by hand), a **Your turn** prompt waits until you pick **Done, continue** or **Cancel**.

The agent needs opencode (`curl -fsSL https://opencode.ai/install | bash`) and a live profile. It uses the active `llms.json` profile, or `code.model: provider/model`. With the `mock` profile, or without opencode, the request runs as one browser goal.

`webpilot run` exits 0 when the job is done, 1 when it failed, 2 when it was stopped, and 3 when the agent needs more information.

## Page memory

Every live run records each page it visits in `.webpilot/memory` next to `webpilot.yaml` (`~/.webpilot/memory` when there is no config file): every interactive element, with all of its locators (test id, id, role and name, label, placeholder, alt, title, href, text, CSS, XPath), a confidence % for each, and when it was first seen, last seen, and last used. Confidence rises when a locator keeps matching exactly one element and works when used, and falls when it goes missing or fails.

A goal that passed before is replayed from memory with no model calls. Each step finds its element by the highest-confidence locator that still matches, so a renamed button or a changed id heals itself. The agent then checks the result with one model call (`replay: verify`), or the run finishes without the model (`replay: trust`). If a step cannot be found, the agent takes over from that point. Generated specs use the best locator from memory.

```bash
webpilot memory                          # hosts, pages, elements, flows
webpilot memory show booking.com         # a host's pages
webpilot memory show https://app.test/login --locators   # every locator with confidence
webpilot memory flows --steps            # recorded goals and their steps
webpilot memory clear booking.com --yes
webpilot run "..." --replay trust       # or --replay off, --no-memory
```

## Tests

Point WebPilot at a repo and a run that passes becomes a test in that repo. The agent gets the run's steps, the locators and confidence from page memory, and the generated spec. It studies the repo's framework, page objects, fixtures, and helpers, reuses them, and writes the test. WebPilot then runs the test command itself and hands any failure back, until the test passes or the attempts run out. The test counts only when WebPilot's own run of it passes.

```bash
curl -fsSL https://opencode.ai/install | bash   # or: npm i -g opencode-ai

webpilot run "Log in and open the invoices page on staging.app.test" --repo ../app
webpilot code --repo ../app          # write a test for the latest run
webpilot code out/20260928-172700 --repo ../app --attempts 3
```

In the shell and `webpilot run`, every passed run becomes a test unless you say otherwise: in `--repo`, in `code.repo` from `webpilot.yaml`, or else in the current folder. `--no-code` turns tests off. The transcript shows the handover to the coding agent, the files it reads and writes, and WebPilot's run of the test, and each request ends with a card listing the run's report, the recorded spec, and the test with its result (or why no test was written).

### Fixing broken tests

When the app changes, tests break. `webpilot heal` runs the repo's tests, and for each failing test the agent re-runs its flow in the real browser. If the flow still works, the test is out of date: the agent updates its steps and locators (in the page object, when that's where they live) from what the browser just did and from page memory. If the browser fails at the same step, the app is broken there: the test is left alone and the failure is reported as an app bug. Tests are never deleted, skipped, loosened, or given longer timeouts to make them pass.

WebPilot re-runs every healed test and then the whole command itself; a test counts as healed only when WebPilot's own run of it passes.

```bash
webpilot heal --repo ../app                              # code.test_command, or the agent finds the command
webpilot heal "npx playwright test tests/checkout" --repo ../app
webpilot heal -- npx playwright test --project=chromium
webpilot run "The login tests broke after the redesign, fix them" --repo ../app   # the agent heals them the same way
```

The report is `heal.md` (and `heal.json`) in `out/heal/<time>/`, with each test's verdict, the files changed, the browser runs, and the test output before and after. `webpilot heal` exits 0 when the tests pass (or nothing needed healing), 1 when some still fail, and 2 when stopped.

## Test cases and suites

Run existing test cases instead of typing goals: ask for them by name ("run SHOP-12", "run test plan 12 suite 34", "run the cases in smoke.yaml") or paste them into the message. Each case becomes one goal: its steps are what the browser agent does and its expected results are what it checks, and the agent's own verdict decides pass or fail. With a repo set, every passed case also gets a test.

`webpilot suite` runs cases the same way every time, with no agent choosing anything, which is what pipelines need. It takes explicit references:

```bash
webpilot cases cases.yaml                     # what a source gives, without running anything
webpilot suite cases.yaml --parallel 3        # run them, 3 browsers at once
webpilot suite ado:plan=12/suite=34 --repo ../app
webpilot suite jira:jql="project = SHOP AND labels = smoke"
webpilot suite xray:plan=SHOP-100
webpilot suite "Open example.com and check the title says Example Domain"
```

| Source | Reference |
|--------|-----------|
| Plain text | the text itself, quoted |
| Text or Markdown file | `.txt` / `.md`, one case per block between `---` lines |
| CSV | `.csv` with `id, title, url, description, preconditions, tags, action, data, expected` columns |
| YAML / JSON | `.yaml` / `.json`, a list of `{id, title, url, preconditions, steps: [{action, data, expected}]}` |
| Gherkin | `.feature`, one case per scenario (and per Examples row) |
| Azure DevOps | `ado:123,456` (test case ids), `ado:plan=12`, `ado:plan=12/suite=34` |
| Jira | `jira:SHOP-1,SHOP-2`, `jira:jql=<query>` (Cloud and Server/Data Center) |
| Xray | `xray:SHOP-1`, `xray:jql=<query>`, `xray:plan=SHOP-100` (Cloud and Server/Data Center) |
| TestRail | `testrail:C12,C13`, `testrail:run=45`, `testrail:plan=7`, `testrail:suite=3` |
| Zephyr Scale (Cloud) | `zephyr:SHOP-T1,SHOP-T2`, `zephyr:cycle=SHOP-R5`, `zephyr:folder=12` |
| qTest | `qtest:1234` (test case ids), `qtest:cycle=88`, `qtest:suite=9` |

For a case with numbered steps the browser agent reports a verdict for every step (passed, failed, or not run, with what it actually saw), not only for the case.

Results go back where the cases came from, with the step verdicts: an Azure DevOps test run for a plan (or a comment on the work item), a Jira comment, an Xray Test Execution with the failure screenshot, a TestRail run (the run the cases came from, or a new one) with step results and the screenshot, a Zephyr Scale test cycle, or qTest test logs. `--no-write-back` or `sources.write_back: false` turns it off.

### Bugs

With `sources.bugs`, each failed case gets a bug in Jira or Azure DevOps with the steps and verdicts, where it ended, the build link, and the last screenshot, linked to the test case. WebPilot tags the bug with a fingerprint of the case, so the next failure adds a "still failing" comment to the open bug instead of filing another, and a pass adds a note that it passes again.

```yaml
sources:
  bugs: { system: jira, project: SHOP }          # or system: azure_devops (uses azure_devops.project)
```

### Reports

Every suite writes `report.html`, `junit.xml`, and `suite.json` to `out/suites/<time>/`, with each case's run folder beside them; every single run writes its own `report.html` too. The report is one HTML file with the screenshots embedded, so you can open it straight from disk, attach it to a build, or mail it; no server is needed. It has an overview (verdict, totals, step checks, tokens, cost, page memory savings, browser errors, accessibility issues, bugs filed, what changed since the last run, results over earlier runs, the cases that need attention, results by source and tag, the slowest cases, and what was posted where), a test results page (each case's step table with expected and actual results, why it failed, its last passing screenshot beside this run's, browser errors per step, page memory with healed locators, the browser agent's steps with a screenshot each and the clicked element outlined, a timing waterfall, accessibility issues, and the trace, spec, and workflow to read in place), a failures page, since last run, trends and flaky tests from the earlier suite runs in the same `out/suites/` folder, time and cost (model vs browser time, a parallel worker timeline, cost per case from the model's `price` in `llms.json`), coverage by requirement and group from the test system, accessibility checks for every page visited, and the environment (model, browser, WebPilot version, CI build, branch, commit). It exports Markdown and JSON and prints cleanly. See `samples/reports/nightly/20260929-020005/report.html` for an example. `report.screenshots: failures` keeps screenshots for failed runs only, `off` drops them; `report.accessibility: false` skips the page checks. On GitHub Actions the suite also writes a job summary.

## Signing in

Sites that need a login go under `logins:`. The browser agent types credentials from the environment without ever seeing them (it only sees placeholders such as `shop_password`), fills authenticator codes from a TOTP secret, and keeps the signed-in session in `.webpilot/sessions/` so the next run starts signed in. Replays and generated Playwright specs read the same environment variables, and specs compute TOTP codes themselves.

```yaml
logins:
  shop:
    url: https://staging.shop.test/login
    username_env: SHOP_USER
    password_env: SHOP_PASSWORD
    totp_env: SHOP_TOTP_SECRET      # optional
```

For single sign-on, captchas, or security keys, sign in by hand once: `webpilot login shop` opens a browser, you sign in, press Enter, and the session is saved. `webpilot logins` lists logins and sessions; `webpilot logins clear shop` forgets one. Session files are private to your user and have their own `.gitignore`.

## Azure DevOps, Jira, and other systems

WebPilot's agent connects to each system's MCP server, configured from the same `sources:` settings in `webpilot.yaml`: organization, project, site URL, and the names of the environment variables that hold the tokens. You don't write any MCP config yourself.

| System | MCP server | Needs |
|--------|------------|-------|
| Azure DevOps Services | Microsoft's [Azure DevOps MCP server](https://github.com/microsoft/azure-devops-mcp) (`npx @azure-devops/mcp`), with work items, boards, test plans, pipelines, and repos | Node.js, and `ADO_PAT` or `az login` |
| Jira Cloud, Server/Data Center | [mcp-atlassian](https://github.com/sooperset/mcp-atlassian) (`uvx`), with issues, JQL search, comments, and transitions. `mcp: rovo` uses [Atlassian's hosted server](https://github.com/atlassian/atlassian-mcp-server) instead (Cloud, and an admin must allow API tokens) | uv; `JIRA_EMAIL` and `JIRA_API_TOKEN` (Cloud) or a personal access token (Server/DC) |
| Confluence | mcp-atlassian, shared with Jira (or its own when Jira isn't set up); Atlassian's hosted server covers it on the same site | `confluence: true` on a Jira Cloud site, or `confluence.url` + tokens |
| Xray, TestRail, Zephyr Scale, qTest | none; WebPilot's API | `XRAY_CLIENT_ID` / `XRAY_CLIENT_SECRET`, `TESTRAIL_EMAIL` / `TESTRAIL_API_KEY`, `ZEPHYR_API_TOKEN`, `QTEST_TOKEN` |
| Anything else | servers you add under `mcp:` | |

```yaml
sources:
  azure_devops: { org_url: https://dev.azure.com/acme, project: Shop }   # token in ADO_PAT
  jira: { url: https://acme.atlassian.net }                              # JIRA_EMAIL + JIRA_API_TOKEN
mcp:
  github:
    url: https://api.githubcopilot.com/mcp/
    headers: { Authorization: "Bearer {env:GITHUB_TOKEN}" }
```

Tokens never go into the config: servers get them through `{env:NAME}` references that opencode fills in, and the agent never sees them. `webpilot mcp list` shows what your settings give. When the agent starts, the shell and `webpilot run` print which servers connected.

Whatever has no MCP server, or whose server didn't connect, goes through WebPilot's own tools, which call the REST APIs directly: `find_test_cases` and `run_test_cases` for test cases and posting results back, `read_confluence` for Confluence pages, and `query_api` to read anything else from Azure DevOps, Jira, Xray, or Confluence. `query_api` only reads (GET requests and Xray GraphQL queries) and only reaches the configured hosts. `sources.azure_devops.mcp: false` or `sources.jira.mcp: false` uses the API only.

Requirements work too: "write and run test cases for the checkout rules page in Confluence" makes the agent read the page, write one case per rule or acceptance criterion (with the main negative paths), show them, and run them. It saves them into a test system only when you ask.

The agent creates or changes items (a bug, a comment, a status change) only when you ask it to; bugs for failed cases come from `sources.bugs`, not from the agent.

## Pipelines

Ask for the pipeline you want and the agent builds it:

```bash
webpilot run "Add a GitHub Actions workflow that runs the smoke cases in cases.yaml on every pull request"
webpilot run "Set up an Azure DevOps pipeline that runs test plan 12 every night and posts results back"
webpilot run "The WebPilot pipeline on main is failing, fix it"
```

It works in the repo (the working directory, or `--repo`): it reads the repo's existing CI, writes or updates the workflow so it installs Bun, uv and WebPilot, calls `webpilot suite <source...> --no-code --plain`, and publishes `junit.xml` and the report, and validates it (actionlint when it is installed). Then it commits to a new branch named `webpilot/ci-<topic>`, stores the keys the pipeline needs as secrets, triggers the run with `gh` or `az`, watches it, and fixes failures until the run passes (up to `code.max_attempts` failed runs).

Guard rails:
- Pushing goes through WebPilot's `push_branch` tool, which refuses the default branch, `main`, and `master`, and never forces. `git push` itself is blocked for the agent.
- Secrets go through `set_ci_secret`: the agent names an environment variable and WebPilot passes its value straight to `gh secret set` or `az pipelines variable`. The agent never sees the value.
- It needs `gh` (GitHub, logged in) or `az` with the `azure-devops` extension (Azure DevOps, logged in). Without them it still writes, validates, and commits the pipeline, then tells you what to run.

## Monitoring

`webpilot monitor` runs checks on a schedule and tells your team when one breaks. A check is any test case `webpilot suite` takes: a YAML file, a Jira key, an Azure DevOps plan. Each round replays the flows from page memory, so a flow that passed before needs one model call to confirm the result (or none with `replay: trust`).

```yaml
monitor:
  checks: [checks/smoke.yaml, "jira:jql=labels = monitor"]
  every: 15m
  fail_after: 2           # two failing rounds in a row before the first alert, to ride out a flaky run
  remind_every: 2h        # repeat while it keeps failing
  alerts:
    teams: { webhook_env: TEAMS_WEBHOOK_URL }
    email: { host: smtp.office365.com, port: 587, to: [qa-team@example.com] }   # SMTP_USERNAME, SMTP_PASSWORD
```

```bash
webpilot monitor --test-alerts      # a sample alert, to check the webhook and SMTP settings
webpilot monitor                    # every 15 minutes until Ctrl+C
webpilot monitor --once             # one round, for cron or a scheduled pipeline
```

An alert goes out when a check starts failing, again every `remind_every` while it fails, and once when it recovers. One round sends one message covering every check that changed. Teams gets an Adaptive Card (both incoming webhooks and Workflows webhooks accept it) and email gets the failing checks' screenshots attached. Each alert links to the round's report: the local file, or `<report_url>/<round>/report.html` when you publish `out/monitor/` somewhere. A round that cannot run at all (a test system is down, the engine will not start) alerts as the check "WebPilot monitor".

Rounds go to `out/monitor/<time>/`, the same as suites, so the report shows trends and flaky checks over the rounds; the last `keep` rounds (100) are kept. `state.json` remembers which checks are failing and since when, so with `--once` in a pipeline, cache that folder between runs. Only one monitor can use a folder at a time. `--once` exits 0 when every check passed and 1 otherwise. Results are posted back to the test system only with `write_back: true`. In the shell, "monitor the smoke cases every 10 minutes" runs the rounds in the background while you keep working, and each round shows as a card; they stop when the shell closes.

## MCP

`webpilot mcp` serves WebPilot's tools over stdio for other agents such as Claude Code, Cursor, or Codex:

```json
{ "mcpServers": { "webpilot": { "command": "webpilot", "args": ["mcp", "--repo", "/path/to/app"] } } }
```

Tools: `browser_run`, `find_test_cases`, `run_test_cases`, `run_details`, `run_test`, `page_memory`, `read_confluence`, `query_api`, `push_branch`, `set_ci_secret`, and one tool per command: `write_test`, `heal_tests`, `monitor`, `logins`, `webpilot_setup`. Browser runs and suites log their progress to stderr, and long calls send progress notifications.
