Metadata-Version: 2.4
Name: ai-infra-bench
Version: 0.1.4.post4
Summary: AI Infra bench
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: aiohttp
Requires-Dist: openai
Requires-Dist: plotly
Requires-Dist: numpy
Requires-Dist: psutil
Requires-Dist: requests
Requires-Dist: PyYAML
Requires-Dist: tqdm
Requires-Dist: hydra-core
Provides-Extra: data
Requires-Dist: datasets; extra == "data"
Requires-Dist: transformers; extra == "data"
Requires-Dist: ai-infra-bench-dataset==0.1.4.post4; extra == "data"
Provides-Extra: deepswe
Requires-Dist: datacurve-pier>=0.3.0; python_version >= "3.12" and extra == "deepswe"
Dynamic: license-file

<div align="center">

![ai_infra_bench](assets/main.png)

[![LICENSE](https://img.shields.io/badge/license-Apache_2.0-orange.svg)](LICENSE)
[![PYTHON VERSION](https://img.shields.io/badge/python-3.10+-blue)](https://www.python.org/)
[![PYPI PROJECT](https://img.shields.io/pypi/v/ai-infra-bench?color=green)](https://pypi.org/project/ai-infra-bench/)

</div>

# AI Infra Bench

AI Infra Bench measures OpenAI-compatible inference endpoints, evaluates model
outputs, and finds the highest load that satisfies a service-level objective.
It is backend-independent and does not require a serving framework SDK.

## Install

```bash
pip install ai-infra-bench
```

Python 3.10 or newer is required.

## CLI

Send one request:

```bash
aib req --base-url http://127.0.0.1:30000 --prompt "Who are you?"
```

Run a random-token benchmark:

```bash
aib bench \
  --base-url http://127.0.0.1:30000 \
  --dataset random \
  --input-len 1024 \
  --output-len 256 \
  --num-requests 100 \
  --max-concurrency 16
```

Other commands cover dataset evaluation, logits and hidden-state comparison,
metric plotting, and local Prometheus monitoring:

```bash
aib --help
```

## SLO Search

SLO searches are configured in YAML. The command probes the configured range
and returns the highest `max_concurrency` or `request_rate` that satisfies every
condition:

```yaml
endpoint:
  base_url: http://127.0.0.1:30000
  api_key: EMPTY
  model: null
request:
  payload:
    messages:
      - role: user
        content: hello
benchmark:
  num_requests: 100
  warmup_requests: 5
  max_concurrency: 32
  request_rate: inf
search:
  parameter: max_concurrency
  min: 1
  max: 64
conditions:
  - metric: success_rate
    operator: ">="
    value: 0.99
  - metric: p99_latency_ms
    operator: "<"
    value: 3000
```

Run it with:

```bash
aib slo examples/slo.yaml -o slo-result.yaml
```

## Output

SLO runs write a YAML report with every probe, its metrics, and condition results.
Request benchmarks can also write machine-readable JSON or JSONL metrics for
later visualization.

## Development

```bash
python -m pip install -e .
pytest
```

Licensed under the [Apache License 2.0](LICENSE).
