Metadata-Version: 2.4
Name: syvain-metrics-collector
Version: 0.0.351
Requires-Dist: typing-extensions>=4.15.0
Requires-Dist: pytest>=8.0.0 ; extra == 'dev'
Requires-Dist: ruff>=0.15.12 ; extra == 'dev'
Requires-Dist: syvain-metrics-api-client>=0.0.334 ; extra == 'dev'
Requires-Dist: ty>=0.0.34 ; extra == 'dev'
Provides-Extra: dev
Summary: Syvain metrics collection SDK
Requires-Python: >=3.11
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM

# Syvain Metrics Collector

Use this Python package to send metrics and annotations from a training,
evaluation, or benchmark job to [Syvain Metrics](https://metrics.syvain.com/).
Use `syvain-metrics-api-client` or `syvain-metrics` cli (available in npm) to read stored metrics.

## Install

```bash
uv add syvain-metrics-collector
```

### Wheels

The package ships compiled wheels for Linux x86_64, Linux aarch64, and macOS
arm64 on CPython 3.11 and later. The queue, batching, retries, and HTTP
delivery run in a Rust core, so a wheel is required; there is no pure Python
fallback.

## Collect an experiment

```python
from syvain_metrics_collector import Collector

collector = Collector(api_key="ak_org_...")
experiment = collector.experiment(
    slug="mamba-run-001",
    description="Baseline mamba training run",
    folder_id="00000000-0000-0000-0000-000000000000",
    meta={
        "model": "mamba",
        "dataset": "internal-v1",
        "seed": 7,
        "config": {"batch_size": 32, "learning_rate": 0.0003},
    },
)

with experiment.run():
    for step in range(1, 1_001):
        # run actual training
        loss = 1.0 / step

        if step == 1 or step % 10 == 0:
            experiment.metric(
                "loss",
                loss,
                step=step,
                metadata={"split": "train"},
            )

    experiment.annotation(
        "Checkpoint saved",
        metadata={"path": "checkpoints/mamba-run-001/step-999.pt"},
    )

experiment.flush_or_raise()
```

`Collector(...)` checks the API key, and the default folder when one is given,
before it returns. Omit `api_key` to read it from the `SYVAIN_METRICS_API_KEY`
environment variable. Other arguments:

| Argument | Default | Meaning |
| --- | --- | --- |
| `host` | `https://metrics.syvain.com` | API base URL |
| `folder_id` | `None` | Folder every new experiment is placed in |
| `max_queue_items` | `100_000` | Queue capacity; over it the oldest metric or annotation is evicted |
| `max_batch_items` | `500` | Events per request |
| `flush_delay_seconds` | `0.25` | How long a burst is coalesced before sending |
| `request_timeout_seconds` | `10.0` | Per-request timeout |
| `logger` | `logging.getLogger("syvain.metrics")` | Logger for dropped values and exit warnings |

`collector.experiment(...)` opens the experiment and blocks until the backend
accepted it, so `experiment.id` and `experiment.url` are available right after
it returns. Its `folder_id` overrides the collector default for that experiment
only. Opening the same slug again returns the same `Experiment` without a
second open.

`metric()` and `annotation()` enqueue data. The collector sends queued batches
in the background. `experiment.run()` records the lifecycle: running on entry,
done on a clean exit, and an error event carrying the exception type and
message when the block raises, before re-raising it.

The final `flush_or_raise()` fails the job if queued data cannot be delivered
or the collector previously evicted events after exceeding `max_queue_items`.
Capacity drop counts persist for the collector's lifetime, including after a
successful queue drain. At process exit the collector drains for up to 30
seconds and logs a warning when events remain pending or were dropped.

Use exactly one `flush_or_raise()` after the `run()` block. Do not call it from
the training loop, evaluation loop, reporting branch, or checkpoint branch.

Annotation text is limited to 16,000 characters. Each annotation's metadata is
limited to 64 KiB, or 65,536 UTF-8 bytes, of compact JSON, including keys, nested
values, and JSON punctuation. The SDK validates annotations before queueing;
the API also rejects oversized metadata. Neither truncates the payload.
Store larger arrays or raw evaluation records as artifacts and put their paths
or URLs in annotation metadata. Existing larger annotations remain readable.

## Collect only measurements the experiment needs

Define the evidence before adding metrics. Each metric must be required to
answer the experiment's question or to interpret training health. Do not emit
every intermediate, tensor statistic, layer value, or runtime diagnostic.

Report training metrics at a planned cadence. Aggregate device tensors first,
then convert them to Python numbers only when reporting. This avoids a device
synchronization on every microbatch.

## Put values in the right field

| Value | Field |
| --- | --- |
| Stable run identity and configuration | Experiment `meta` |
| Numeric measurement | Metric `value` |
| Training or evaluation progress | Metric `step` |
| Bounded category used to group a series | Metric `metadata` |
| Unique event details, paths, hashes, IDs, and text | Annotation `metadata` |

Use one stable metric name for one quantity and unit. Keep the same name across
splits, datasets, stages, devices, and ranks:

```python
experiment.metric("loss", train_loss, step=step, metadata={"split": "train"})
experiment.metric("loss", valid_loss, step=step, metadata={"split": "valid"})
```

Do not encode dimensions in the name:

```python
# Wrong
experiment.metric(f"{stage}/{split}/loss", loss, step=step)

# Correct
experiment.metric(
    "loss",
    loss,
    step=step,
    metadata={"stage": stage, "split": split},
)
```

## Keep metric metadata low-cardinality

Metric metadata is a flat `str -> str` mapping. Use it only for bounded
categories needed to compare series, such as `split`, `dataset`,
`training_stage`, `device`, `rank`, or optimizer parameter group.

Every distinct metadata mapping creates a separate series. The product of all
dimension values, including missing-key variants, must stay at or below 4,096
series per metric in one experiment. For example, 8 stages, 3 splits, and 16
ranks produce 384 series.

Never put steps, epochs, timestamps, paths, sample or request IDs, hashes, free
text, numeric measurements, or serialized objects in metric metadata. Put
progress in `step`, stable configuration in experiment `meta`, and unique
details in annotations.

The client accepts at most 32 metadata keys, 128 UTF-8 bytes per key, 512 UTF-8
bytes per value, and 4,096 UTF-8 bytes in the canonical JSON mapping. It
validates these limits before enqueueing the metric and raises `ValueError`
when one is exceeded. Metric values must be finite numbers; the client logs and
drops non-finite values. Pass `timestamp=` seconds since the epoch, or
milliseconds at that scale, to stamp a measurement yourself instead of at
enqueue time.

## Use test collectors

Use `NoopCollector()` when a test only needs the collector interface; it does
no IO and reports every event as delivered. Use `JsonlCollector(path=...)` when
a local run needs inspectable output; it appends one JSON object per event to
the file and needs no API key. Both expose the same `experiment()`, metric,
annotation, lifecycle, and flush calls as `Collector`.

```python
from syvain_metrics_collector import JsonlCollector, NoopCollector

silent = NoopCollector()
local = JsonlCollector("metrics.jsonl")
```

