Metadata-Version: 2.5
Name: twinbox
Version: 0.1.0
Summary: Black-box parity testing of two implementations of an HTTP service, run in Docker
Project-URL: Homepage, https://gitlab.com/twinbox-group/twinbox
Project-URL: Repository, https://gitlab.com/twinbox-group/twinbox.git
Project-URL: Issues, https://gitlab.com/twinbox-group/twinbox/-/issues
Author: Mikhail Lopotkov
License-Expression: MIT
License-File: LICENSE
Keywords: black-box,docker,http,parity,pytest,testing
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Framework :: Pytest
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Testing
Classifier: Typing :: Typed
Requires-Python: >=3.13
Requires-Dist: aiohttp>=3.14.3
Requires-Dist: click>=8.2
Requires-Dist: docker>=7.2
Requires-Dist: httpx>=0.28.1
Requires-Dist: locust>=2.46
Requires-Dist: psycopg[binary]>=3.3.6
Requires-Dist: pydantic>=2
Requires-Dist: pytest
Requires-Dist: pytest-html>=4.2
Requires-Dist: pytest-xdist>=3.8.0
Requires-Dist: pyyaml
Requires-Dist: requests>=2.34.2
Requires-Dist: testcontainers>=4.15.0
Description-Content-Type: text/markdown

# twinbox

<div align="center">
  <img src="docs/logo.png" alt="twinbox logo" width="240">
</div>

**Check that a rewritten HTTP service behaves exactly like the original.**

[Русская версия](README.ru.md)

twinbox is a Python library for **black-box parity testing of HTTP services**.
It is meant for teams that rewrite a service — for example, from Python to Go
or Rust — and need to show that the new implementation behaves the same way
from the outside as the old one.

You describe the expected behaviour once, as a corpus of YAML scenarios
(request → expected response). twinbox starts each implementation in a Docker
container on a fresh, isolated environment, runs the same scenarios against
it and compares the live responses with the expectations. Differences you
accept on purpose are written into the scenario together with the reason, so
they stay visible instead of being silently ignored.

> **Status: early development.** There is no release yet, and the API and
> the scenario format may still change.

## What works today

- **Corpus check without containers** — `twinbox check` validates every
  scenario (structure, tags, masks, substitutions, accepted differences,
  coverage matrix) and reports each problem with the field path and an
  error code.
- **Run against one implementation** — `twinbox run <name>` starts a
  PostgreSQL database, a mock of external services and the implementation
  itself, runs the scenarios and writes a JSON result document and,
  optionally, an HTML report.
- **pytest integration** — every scenario is an ordinary pytest test, so
  `-k`, `-m` and parallel runs with `pytest-xdist` work as usual.
- **Load comparison** — `twinbox load` runs locust against each
  implementation in turn, with the same CPU and memory limits, and writes one
  Markdown report: latency and RPS by handle, CPU and memory of the container,
  CPU cost of single calls, and a verdict by handle from the project's own
  criterion.

Planned after the first release: code coverage of the reference
implementation, through integration plugins.

## How it works

For every scenario twinbox builds a clean environment:

1. a fresh copy of a PostgreSQL database with the schema applied;
2. a mock of the external services the implementation calls — it echoes
   requests by default and can be programmed per scenario;
3. the implementation under test, started from its Docker image.

Then it sends the scenario's requests, compares each response with the
expectation and tears the environment down. Nothing leaks from one scenario
into the next.

## Requirements

- Python 3.13+
- [uv](https://docs.astral.sh/uv/)
- Docker; for `twinbox load` also cgroup v2 on the host and `cat` in the
  implementation's image (see [docs/usage.md](docs/usage.md#environment))

## Installation

Until the first release, install from the repository:

```sh
uv add git+https://gitlab.com/twinbox-group/twinbox.git
```

## Quick start

**1. Enable twinbox in your project's `pyproject.toml`.** Every field has a
default; set only what differs:

```toml
[tool.twinbox]
plugin = "project_plugin"   # module that describes your implementations
scenarios_dir = "scenarios" # where scenario.yaml files live
```

**2. Describe your implementations** in that module — one `Profile` per
implementation, exactly one of them marked as the reference:

```python
from twinbox.harness.profile import Endpoint, Profile, ReadinessProbe

_API = (Endpoint(name="api", port=8080, default=True),)
_READY = ReadinessProbe(path="/health", expected_status=200)

profiles = [
    Profile(
        name="legacy",
        is_reference=True,
        default_image="shop/legacy:dev",
        named_endpoints=_API,
        readiness_probe=_READY,
    ),
    Profile(
        name="rewrite",
        default_image="shop/rewrite:dev",
        named_endpoints=_API,
        readiness_probe=_READY,
    ),
]
```

**3. Write a scenario**, e.g. `scenarios/users/create_and_fetch/scenario.yaml`:

```yaml
description: create a user and fetch it by the returned id
tags: []  # tags must be declared in the plugin's `labels`; none are used here
steps:
  - name: create user
    request:
      method: POST
      path: /users
      json: {name: Ann}
    expect:
      status: 201
      json_exact:
        id: "{{any_uuid}}"            # mask: any UUID matches
        created_at: "{{any_iso8601}}" # mask: any ISO 8601 timestamp matches
        name: Ann
    capture:
      - {name: user_id, path: id}
  - name: fetch user by captured id
    request:
      method: GET
      path: "/users/{{capture.user_id}}"  # value captured in the previous step
    expect:
      status: 200
      json_subset: {id: "{{capture.user_id}}"}
```

**4. Check and run:**

```sh
uv run twinbox check                                           # validate the scenarios
uv run twinbox run legacy --image shop/legacy:dev --html reports/legacy.html
uv run twinbox run rewrite --image shop/rewrite:dev --html reports/rewrite.html
```

Run `twinbox help` or `twinbox help <command>` for the full reference of
every command and option.

Results are written to `artifacts/results/` (JSON) and, when `--html <path>`
is passed, to the specified file path (HTML).

## Load comparison

`twinbox load` answers the question "is the rewrite more expensive under
load?". For each implementation, one after another, it starts a fresh stand
with the CPU and memory limits from the profile, calls the project's
`load_dataset` to create data, waits for the CPU to calm down and warms the
implementation up. Then it measures: a locust ladder with the read locust
file, the same with the write locust file, and the CPU cost of single calls
to chosen handles. The ladder adds users step by step until the container
reaches its CPU limit, requests start to fail, or the steps run out; in the
last case the report says `limit not reached in <max_steps> steps`. The
peak number of database connections is a column of the report and does not
stop the ladder: the connection pool size is set inside the service, and
twinbox does not know it. A service that hits its pool shows it in this
column together with a growing p95.

**1. Give every profile CPU and memory limits:**

```python
from twinbox.harness.profile import ResourceLimits

_LIMITS = ResourceLimits(cpu_nanocores=1_000_000_000, memory_bytes=512 * 1024 * 1024)
# in every Profile(...): resource_limits=_LIMITS
```

**2. Point twinbox at the locust files** in `pyproject.toml`. Every other
field of `[tool.twinbox.load]` has a default:

```toml
[tool.twinbox.load]
read_locustfile = "load/read.py"
write_locustfile = "load/write.py"  # optional
step_duration = "20s"               # default: 30s
max_steps = 10                      # default: 20
```

Settings of locust itself go to `[tool.locust]` or after `--`. Users, spawn
rate, run time, host and report files are set by twinbox on every step, so
`-u`, `-r`, `-t`, `-H`, `-f`, `--csv`, `--html` and `--headless` are rejected
there.

**3. Write the locust file.** `twinbox.load.binding()` returns the data that
`load_dataset` created on the current stand, `twinbox.load.endpoints()` — the
addresses of the implementation's endpoints:

```python
from locust import HttpUser, task

from twinbox.load import binding


class ReadUser(HttpUser):
    def on_start(self) -> None:
        self.user_id = binding()["user_id"]

    @task
    def get_user(self) -> None:
        self.client.get(f"/users/{self.user_id}", name="/users/[id]")
```

**4. Add two functions to the plugin module:** `load_dataset` creates the
data, `acceptance_criterion` decides what is acceptable. The numbers come as
dataclasses from `twinbox.load`:

```python
from collections.abc import Mapping, Sequence

import httpx

from twinbox.load import HandleVerdict, ImplementationNumbers, NetworkNumbers


def load_dataset(endpoints: Mapping[str, str], profile: str) -> Mapping[str, str]:
    response = httpx.post(f"{endpoints['api']}/users", json={"name": profile}, timeout=10)
    response.raise_for_status()
    return {"user_id": response.json()["id"]}


def acceptance_criterion(numbers: Sequence[ImplementationNumbers]) -> Sequence[HandleVerdict]:
    reference = next(item for item in numbers if item.is_reference)
    if not isinstance(reference.read, NetworkNumbers):
        return []  # a failed measurement already makes the exit code 2
    return [
        HandleVerdict(
            handle=handle,
            implementation=item.implementation,
            accepted=own.p95_ms <= reference.read.handles[handle].p95_ms * 1.2,
            note=f"p95 {own.p95_ms:.1f} ms",
        )
        for item in numbers
        if not item.is_reference and isinstance(item.read, NetworkNumbers)
        for handle, own in item.read.handles.items()
        if handle in reference.read.handles
    ]
```

**5. Run:**

```sh
uv run twinbox load                        # every implementation, then the verdict
uv run twinbox load rewrite                # one implementation, numbers only
uv run twinbox load -- --only-summary      # everything after -- goes to locust
```

The report is written to `artifacts/load/<time>/summary.md`, next to
locust's HTML and CSV files of every step; its path is printed at the end.
Exit code `0` means every measurement ran and no verdict is rejected, `1`
means at least one verdict is rejected, `2` means a check before the start or
a measurement failed, the criterion failed, or the report was not written.
Errors are printed as `[<code>] <message>`. While the run goes, lines like
`[load] rewrite read: step 2: starting, 20 users for 30s` on stderr show which
implementation, kind and step is running now.

The CPU and memory limits apply to `twinbox run` too. The full reference —
every setting and its default, the isolated handles, the forms of the
numbers, the report and the errors — is in
[docs/usage.md](docs/usage.md#twinbox-load).

## Examples

A runnable example lives in its own repository,
[gitlab.com/twinbox-group/example](https://gitlab.com/twinbox-group/example):
one calculator service in two implementations, Python and Rust, and nine
lessons that introduce twinbox's capabilities one at a time, starting from a
plain check.

## Documentation

- [docs/usage.md](docs/usage.md) — settings, extension points, the scenario
  format, commands and their results, the load comparison.
- [CONTRIBUTING.md](CONTRIBUTING.md) — development setup, checks, the Docker
  test suite and supported platforms.

## License

[MIT](LICENSE).
