Metadata-Version: 2.4
Name: benchcore
Version: 0.6.0
Summary: A strongly typed, framework-agnostic benchmarking engine for Python
License-Expression: MIT
License-File: LICENSE.md
Keywords: benchmark,performance,testing,typing
Author: Eli-ezer Reuven Ramirez Ruiz
Author-email: ramirez.ruiz.eliezer.reuven@gmail.com
Requires-Python: >=3.12
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Typing :: Typed
Project-URL: Documentation, https://github.com/ezer-mackenzie/benchcore#readme
Project-URL: Issues, https://github.com/ezer-mackenzie/benchcore/issues
Project-URL: Repository, https://github.com/ezer-mackenzie/benchcore
Description-Content-Type: text/markdown

# BenchCore

> [!IMPORTANT]
> BenchCore is pre-alpha. A minimal benchmarking API exists for experimentation,
> but it is not ready for production use and may change without notice.

BenchCore is a strongly typed, framework-agnostic benchmarking engine
for Python. Its goal is to provide the measurement primitives needed by test
framework integrations without making the core depend on any test runner.

The project grew out of a desire for a benchmarking codebase where typing,
maintainability, and ease of contribution are foundational constraints. Existing
tools such as `pytest-benchmark` helped demonstrate the value of benchmarking in
developer workflows. BenchCore is an independent implementation with a broader
architectural goal: a reusable engine that can eventually serve `pytest`,
`unittest`, and other integrations.

BenchCore is **not** a testing framework, a test runner, or a drop-in replacement
for `pytest-benchmark`. It will measure code; the surrounding framework will
remain responsible for discovering and running tests.

## Design goals

- A small, stable, and strongly typed public API.
- Strict Pyright compatibility without exposing `Any` in public interfaces.
- No runtime dependency on `pytest` or another testing framework.
- Replaceable timers, calibration, statistics, reporting, and storage components.
- Near-zero runtime dependencies and low measurement overhead.
- An approachable codebase designed for long-term maintenance.
- Support for Python 3.12 and newer.

These are design constraints, not claims about features already implemented.

## Intended experience

The eventual standalone API should be simple enough to look like this:

```python
from benchcore import Bench

bench = Bench()
result = bench.run(sorted, [5, 4, 3, 2, 1])

print(result.statistics.mean_ns)
```

Or, for the common case:

```python
from benchcore import benchmark

result = benchmark(sorted, [5, 4, 3, 2, 1])
```

Both examples are implemented by the current MVP. Configuration is available
through `Bench`:

```python
from benchcore import Bench, BenchmarkConfig

bench = Bench(BenchmarkConfig(rounds=20, warmup_rounds=2))
result = bench.run(sorted, [5, 4, 3, 2, 1])
```

By default, BenchCore calibrates the iteration count until a round reaches a
target duration. Calibration samples are discarded before warmup and measurement:

```python
config = BenchmarkConfig(
    target_round_time_ns=10_000_000,
    max_iterations=1_000_000,
    max_calibration_time_ns=1_000_000_000,
)
result = Bench(config).run(sorted, [5, 4, 3, 2, 1])

print(result.iterations)
```

Set `iterations` explicitly to disable calibration when a fixed count is required:

```python
config = BenchmarkConfig(iterations=100)
```

`max_calibration_time_ns` is an accumulated calibration budget, not a hard
timeout. BenchCore cannot interrupt a running callable, so one slow invocation may
exceed the budget before calibration stops.

Fixed iterations are preferable when repeated calls change workload cost, mutate
shared state, consume a finite input, or when reproducing an earlier run with an
exact count. Automatic calibration is intended for callables whose cost remains
reasonably stable across repeated invocations.

## Round lifecycle

Optional hooks can prepare and restore state around every calibration, warmup, and
measured round:

```python
values = [3, 2, 1]


def reset_values() -> None:
    values[:] = [3, 2, 1]


bench = Bench(
    BenchmarkConfig(iterations=1),
    round_setup=reset_values,
    round_teardown=values.clear,
)
result = bench.run(values.sort)
```

`round_setup` and `round_teardown` are outside the timed interval. Once setup
succeeds, teardown runs exactly once even when the benchmark callable or timer
fails. If setup itself fails, teardown does not run. A teardown exception
propagates and no partial result is returned.

BenchCore does not modify garbage collection or serialize concurrent runs. Hooks,
callables, arguments, and injected timers must provide any thread safety and
process-global state restoration they require.

## Running the example

Install the development environment and execute the standalone example:

```console
poetry install
poetry run python examples/basic.py
```

The reported durations are normalized per iteration and stored in nanoseconds.
The example uses `format_duration_ns()` to select a readable display unit without
changing the stored numeric data. Benchmark values vary between machines and even
between runs on the same machine; compare results only under controlled
conditions.

## Results and units

Duration-bearing names include their unit explicitly. Whole measured regions use
integer nanoseconds; normalized durations and statistics use floats because one
iteration may represent a fraction of a timer tick:

```python
from benchcore import format_duration_ns

print(result.total_time_ns)
print(result.statistics.mean_ns)
print(format_duration_ns(result.statistics.mean_ns))
```

`standard_deviation_ns` is the sample standard deviation across measured rounds.
It uses Bessel's correction and is `0.0` when only one round exists. These
descriptive statistics summarize observed runtime noise; they do not establish
statistical significance or prove that one implementation is faster.

Quartiles, percentiles, and outlier labels are intentionally deferred until
BenchCore defines minimum sample sizes and interpolation policies.

## Reports and JSON

Reporting is explicit and occurs after measurement. A report excludes the
arbitrary callable return value while retaining measured rounds, statistics, and
minimal environment identity:

```python
from benchcore import BenchmarkReport, JsonReporter, TerminalReporter

report = BenchmarkReport.from_result("sorted-list", result)

print(TerminalReporter().render(report))
json_payload = JsonReporter(indent=2).render(report)
```

Reporters return strings and never print or write files themselves. The canonical
JSON representation can also be used directly:

```python
from benchcore import deserialize_report, serialize_report

payload = serialize_report(report)
restored = deserialize_report(payload)

assert restored == report
```

The current schema version is `1`. Deserialization rejects malformed JSON,
missing or unknown fields, invalid numeric values, inconsistent statistics, and
unsupported schema versions. Schema v1 compatibility is protected by a committed
round-trip fixture.

## Baselines and regression comparison

`JsonFileStorage` persists reports atomically using deterministic SHA-256
filenames derived from their benchmark names:

```python
from pathlib import Path

from benchcore import JsonFileStorage

storage = JsonFileStorage(Path(".benchcore"))
storage.save(report)
baseline = storage.load(report.name)
```

The exact report name is its canonical identity. Include parameters when they
distinguish benchmark cases, for example `sort[size=1000,order=random]`.

Compatible reports can be compared using absolute and relative tolerances:

```python
from benchcore import (
    ComparisonTerminalReporter,
    RegressionThreshold,
    compare_reports,
)

comparison = compare_reports(
    baseline,
    current,
    threshold=RegressionThreshold(
        relative=0.05,
        absolute_ns=1_000,
    ),
)

print(ComparisonTerminalReporter().render(comparison))
```

The effective tolerance is the greater of the relative threshold and the
absolute nanosecond threshold. Current performance is classified as
`improvement`, `stable`, or `regression`. Reports with different names, Python
versions, implementations, or platform identities are rejected rather than
silently compared.

This classification is a practical threshold over mean duration; it is not a
statistical significance test. `ComparisonJsonReporter` provides the same result
as machine-readable JSON.

Run the complete storage and comparison example with:

```console
poetry run python examples/comparison.py
```

## Scope

The core is expected to grow around a small number of benchmarking concepts:

- precise and replaceable timers;
- warmup and calibration;
- iterations and rounds;
- immutable results and descriptive statistics;
- reporters and result storage;
- regression comparison;
- isolated adapters for testing frameworks.

Features will be designed only when there is a concrete use case. BenchCore will
prefer composition and small protocols over speculative abstraction.

## Project status

BenchCore is currently **pre-alpha**. The development version supports explicit
rounds, automatic or fixed iterations, warmups, an injectable nanosecond timer,
and basic descriptive statistics. Async callables, reporters, storage, framework
integrations, and stability guarantees have not been implemented.

No compatibility guarantees apply until an initial public release. Once public
APIs exist, changes will be documented in [CHANGELOG.md](CHANGELOG.md).

## Contributing

Early contributions are especially valuable when they clarify use cases and API
constraints. Before implementing a substantial feature, please open an issue so
the design and trade-offs can be discussed first. See
[CONTRIBUTING.md](CONTRIBUTING.md) for the complete process.

By participating, you agree to follow the
[Code of Conduct](CODE_OF_CONDUCT.md). Please report security issues privately as
described in [SECURITY.md](SECURITY.md).

## License

BenchCore is distributed under the terms in [LICENSE.md](LICENSE.md).

