Metadata-Version: 2.4
Name: critiqor
Version: 0.2.19
Summary: Runtime intelligence CLI for reviewing AI agent reliability.
Author: Critiqor Contributors
License-Expression: MIT
Keywords: ai,agents,openclaw,observability,reliability,diagnosis
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click>=8.1
Requires-Dist: cryptography>=43
Requires-Dist: websocket-client>=1.8
Dynamic: license-file

<p align="center">
  <img src="assets/Critiqor.png" alt="Critiqor logo" width="120" />
</p>

<h1 align="center">Critiqor</h1>

<p align="center">
  <strong>Runtime Intelligence for AI Agents</strong>
</p>

<p align="center">
  Observe. Diagnose. Improve.
</p>

<p align="center">
  <a href="https://pypi.org/project/critiqor/"><img alt="PyPI" src="https://img.shields.io/pypi/v/critiqor?color=20d6ad"></a>
  <img alt="Python" src="https://img.shields.io/badge/python-3.10%2B-blue">
  <img alt="Status" src="https://img.shields.io/badge/status-alpha-orange">
  <img alt="License" src="https://img.shields.io/badge/license-MIT-green">
</p>

<p align="center">
  <code>pip install critiqor</code>
</p>

Critiqor helps developers understand whether an AI agent run can be trusted.
It observes the runtime, preserves evidence, generates an evidence-backed
diagnosis, and opens a local dashboard with a concrete improvement path.

Instead of judging only the final answer, Critiqor looks at what happened while
the agent worked: framework lifecycle events, tool activity, memory behavior,
errors, confidence signals, and whether the next run improved.

This repository is that existing product, plus a WebMCP evaluation layer added
for the [WebMCP Hackathon](https://webmcp.devpost.com/).

![Critiqor dashboard overview](assets/screenshots/dashboard-overview.png)

---

## WebMCP Hackathon

**Live judge experience:** [Open the Critiqor × Crema experiment](https://critiqor-crema-reliability.terrence-qiu-7311.chatgpt.site/?agent-playback=1)

No login, local service, private path, API key, or checkout is required. In a
WebMCP-capable browser, ask the agent: **“Show the improved run, then explain
why it is safer.”** The page exposes five typed WebMCP tools for inspecting the
verified experiment, reading its playbook and method, and visibly replaying
either experiment arm. Humans can inspect the same evidence, genuine Crema
target, and Critiqor dashboards side by side.

**Research question:** Which mechanisms that helped MCP become more production-ready can be adapted to improve WebMCP reliability?

Critiqor already answered whether an agent run can be trusted, why, and whether a later run improved. The hackathon work extends that same observe → diagnose → improve loop to WebMCP, instead of shipping a separate product.

When a run includes WebMCP events, Critiqor now:

- records consequential tool calls, outcomes, and authoritative application state
- treats a lost or timed-out response as `unknown`, not as a safe failure
- flags a blind retry of the same intent before the first outcome is reconciled
- writes a run-specific improvement playbook from that evidence
- compares a matched later run and reports whether the same failure recurred or was resolved

The first implementation focuses on one failure: retrying a consequential WebMCP action after an ambiguous outcome. A raw agent can duplicate an effect. After the playbook, the same task reconciles first and stops at one effect.

### Work added during the submission period

The underlying Critiqor product predates the challenge. The challenge-specific
work, added from August 25 through September 3, 2026, is the WebMCP browser
monitor, normalized runtime evidence, controlled lost-response fault,
reconciliation checks, matched Crema experiment, public anonymized dashboards,
five page tools, and the judge-facing experiment site. The implementation and
reproduction instructions live in
[`explorations/webmcp-reliability`](explorations/webmcp-reliability), with the
runtime adapter documented in
[`docs/webmcp-browser-monitor.md`](docs/webmcp-browser-monitor.md).

---

## Why Runtime Evaluation Matters

An agent can produce a useful-looking response while still behaving unreliably
during execution. It might ignore relevant memory, miss a tool failure, recover
from an error in a way that hides risk, or appear confident without enough
supporting evidence.

Critiqor gives developers a practical review layer for answering:

- Can I trust this agent run?
- Why?
- What evidence supports that diagnosis?
- What should I change?
- Did the improvement work on later runs?

---

## Supported Agent Frameworks

Critiqor 0.2.19 supports framework-based monitoring for:

- OpenClaw
- Claude Code
- Codex CLI
- Custom CLI frameworks configured with `critiqor agents` or `critiqor config`

Critiqor integrates into your existing workflow. It launches or observes the
agent command, lets you work normally, then finalizes the run into a local
diagnosis dashboard.

---

## Installation

Install Critiqor from PyPI:

```bash
pip install critiqor
```

Check the CLI:

```bash
critiqor help
```

Use Python 3.10 or newer. `pipx install critiqor` is a good option if you prefer
an isolated CLI install.

---

## Quick Start

### 1. Choose an agent framework

```bash
critiqor agents
```

The guided setup lets you choose OpenClaw, Claude Code, Codex, or a custom CLI
framework and observation method. For a custom framework, enter the same launch
command you normally use to open its TUI. Escape moves back through setup; when
you cancel, the draft is discarded without changing your saved configuration.

### 2. Start an observation

Use the monitor command for your framework:

```bash
critiqor monitor openclaw
critiqor monitor cc
critiqor monitor codex
critiqor monitor custom my-agent
critiqor monitor webmcp --help
```

Custom frameworks can be launched through the command you configure in the
guided setup. Critiqor starts the real agent command inside its runtime wrapper,
so the normal agent TUI remains available while evidence collection runs in the
background.

If you selected **Import Log**, stage an exported JSON or JSONL log instead:

```bash
critiqor import-log codex
# Finder opens on macOS; after validation:
critiqor finalize
```

Use `--file ./agent-log.jsonl` for automation or terminals without a graphical
file picker. The VS Code / Cursor Extension option is marked **Coming soon** and
cannot be saved as an active observation method yet.

### 3. Work normally

Use the agent as you usually would. Critiqor stays beside the workflow and
collects runtime evidence for review.

### 4. Finalize the run

```bash
critiqor finalize
```

Critiqor stops the observation, generates a diagnosis, and opens the local
dashboard.

### 5. Reopen reports

```bash
critiqor runs
critiqor dashboard
critiqor dashboard run_001
```

---

## CLI Workflow

```text
critiqor agents
        ↓
Select Framework
        ↓
Choose Observation Method
        ↓
Launch Agent
        ↓
Work Normally
        ↓
critiqor finalize
        ↓
Dashboard Opens
```

Core commands:

- `critiqor agents` - choose and configure an AI agent framework
- `critiqor config` - update observation method or custom framework details
- `critiqor import-log <framework>` - select and stage an exported runtime log
- `critiqor monitor openclaw` - launch OpenClaw and begin runtime observation
- `critiqor monitor cc` - launch Claude Code and begin runtime observation
- `critiqor monitor codex` - launch Codex CLI and begin runtime observation
- `critiqor monitor custom <framework>` - launch a configured custom agent
- `critiqor monitor webmcp` - observe live WebMCP activity in one Chrome tab
- `critiqor finalize` - stop observation, generate diagnosis, and open dashboard
- `critiqor dashboard [run_id]` - open the latest or selected diagnosis dashboard
- `critiqor runs` - list completed evaluations with summaries
- `critiqor doctor` - check every configured agent and local runtime dependency

---

## Dashboard

After finalization, Critiqor opens a local dashboard focused on the developer
questions that matter after an agent run.

Key sections:

- **Overview** - production verdict, trust score, confidence, current run, and the
  fastest path to diagnosis, evidence, playbook, and comparison.
- **Runs** - completed evaluations you can reopen and compare.
- **Diagnosis** - the primary issue, root cause, evidence, runtime impact, and
  engineering explanation.
- **Playbook** - recommended changes, verification steps, expected improvement,
  trade-offs, and alternatives.
- **Evidence Explorer** - timeline events, tool calls, memory events, evidence
  status, and raw event snapshots.
- **Visibility** - private, shared, anonymous, and public review modes.
- **Appearance** - readable dashboard display settings.
- **Export Diagnosis** - PDF, Markdown, HTML, PNG, diagnosis JSON, session JSON,
  and ZIP export options.
- **Copy Fix Prompt** - a run-specific prompt you can paste into an AI coding
  assistant to improve the agent using the observed evidence.

The dashboard supports light and dark appearance modes, so exported screenshots
and team reviews can match the environment where developers are working.

![Critiqor Evidence Explorer](assets/screenshots/dashboard-evidence-explorer.png)

---

## Live WebMCP Browser Monitoring

Critiqor 0.2.19 can observe browser-native WebMCP discovery, invocation, outcome,
reconciliation, and authoritative-state events in real time through Chrome's
remote-debugging endpoint. It does not infer a failed consequential action was
uncommitted: an opaque error, cancellation, or intentionally lost response is
recorded as `unknown` until target-owned state reconciles it.

Start a separate Chrome profile with remote debugging enabled. For example, on
macOS:

```bash
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome \
  --remote-debugging-port=9222 \
  --user-data-dir=/tmp/critiqor-chrome
```

Open the WebMCP site in that Chrome instance, then attach Critiqor with the
endpoint stated explicitly:

```bash
critiqor monitor webmcp \
  --cdp-url http://127.0.0.1:9222 \
  --target-url http://127.0.0.1:3000 \
  --task-id add-one-item \
  --scenario-id lost-response \
  --consequential-tool add_to_cart \
  --reconciliation-tool get_cart \
  --authoritative-tool get_cart
```

`CRITIQOR_CDP_URL` can supply the endpoint instead of `--cdp-url`. The target
URL must match exactly one open page by URL prefix. Critiqor does not enable
remote debugging in an arbitrary Chrome process; Chrome must expose or approve
the endpoint first. Press Ctrl-C after the browser task, then run `critiqor
finalize` to generate the diagnosis and playbook.

Fault injection is optional and deliberately narrow. For the Crema cart
mutation experiment, bind the one-shot response fault to the exact URL, HTTP
method, and consequential WebMCP tool:

```bash
critiqor monitor webmcp \
  --cdp-url http://127.0.0.1:9222 \
  --target-url http://127.0.0.1:3000 \
  --task-id add-one-bianca \
  --scenario-id commit-lost-response \
  --consequential-tool add_to_cart \
  --reconciliation-tool get_cart \
  --authoritative-tool get_cart \
  --allowed-api-origin http://localhost:3001 \
  --fault-response-url http://localhost:3001/operations/add-to-cart \
  --fault-method POST \
  --fault-tool add_to_cart \
  --authoritative-state-url http://localhost:3001/operations/get-cart
```

The adapter injects at most once and only when exactly one matching tool
invocation is pending. It refuses an ambiguous concurrent correlation. See
`docs/webmcp-browser-monitor.md` for the setup and evidence contract.

## What's New in 0.2.19

Critiqor 0.2.19 adds a public Chrome/WebMCP runtime adapter.

- `critiqor monitor webmcp` connects to an explicit Chrome remote-debugging
  endpoint and validates WebMCP/CDP support before creating a run.
- Live registry, dispatch, outcome, reconciliation, and authoritative-effect
  evidence is normalized into Critiqor's WebMCP event vocabulary.
- Optional response-stage fault injection supports controlled lost-response
  experiments without claiming that an ambiguous action failed safely.
- `websocket-client` is now installed as a runtime dependency.

## What's New in 0.2.18

Critiqor 0.2.18 adds WebMCP runtime evaluation when a run includes WebMCP
events, plus a tighter dashboard review path.

- WebMCP runs produce an evidence-backed diagnosis, a run-specific improvement
  playbook, and a detailed Copy Fix Prompt from the selected run artifacts.
- Diagnosis, Playbook, and Evidence share a Focus run dropdown bound to
  `run_id`, so another run is never substituted.
- Engineer Brief, Executive Summary, and Agent Health cards open the same
  keyboard-accessible detail view. Missing fields stay unavailable.
- The local dashboard is served from the bundled production build.

## What's New in 0.2.16

Critiqor 0.2.16 focuses on runtime memory evaluation and the matching dashboard
experience.

- Memory behavior is included in the diagnosis workflow when evidence is
  available.
- Retrieved, injected, referenced, unused, irrelevant, missed, created, ignored,
  and not-stored memory events can be explained from runtime evidence.
- Copy Fix Prompt includes memory behavior, supporting evidence, suggested
  architectural improvements, testing strategy, and success criteria.
- The dashboard reflects the current diagnosis, evidence, playbook, export, and
  visibility workflow.
- OpenClaw, Claude Code, Codex CLI, and custom framework workflows are presented
  as first-class ways to observe AI agents.

---

## Export and Team Review

Critiqor reports can be used to:

- improve prompts, tools, memory, and agent architecture
- share a diagnosis with teammates
- document runtime evaluations
- compare whether changes improved later runs
- provide evidence for release or review decisions

Export options include PDF, Markdown, HTML, PNG, diagnosis JSON, session JSON,
and ZIP bundles.

---

## Visibility Modes

Critiqor supports dashboard visibility settings from the developer's point of
view:

- **Private** - local owner review.
- **Shared** - invite-based review for teammates.
- **Anonymous** - redacted review without exposing identifying details.
- **Public** - open dashboard access when you intentionally choose it.

Configure visibility through `critiqor config`, then relaunch the dashboard.

---

## Operating System Compatibility

| Operating system | Compatibility | Recommended install path |
| --- | --- | --- |
| macOS | Supported | Python 3.10+ with `pip` or `pipx` |
| Linux | Supported | Distro Python package manager, then `pip` or `pipx` |
| Windows | Supported with WSL recommended | WSL2 for terminal agent workflows, or native Windows Python for basic CLI usage |

For the most reliable terminal-agent monitoring on Windows, use WSL2.

---

## License

MIT
