Metadata-Version: 2.3
Name: syvain-metrics-collector
Version: 0.0.331
Summary: Syvain metrics collection SDK
Requires-Dist: niquests>=3.16.0
Requires-Dist: syvain-metrics-api-client>=0.0.331
Requires-Dist: typing-extensions>=4.15.0
Requires-Dist: pytest>=8.0.0 ; extra == 'dev'
Requires-Dist: ruff>=0.15.12 ; extra == 'dev'
Requires-Dist: ty>=0.0.34 ; extra == 'dev'
Requires-Python: >=3.10
Provides-Extra: dev
Description-Content-Type: text/markdown

# Syvain Metrics Collector

Use this Python package to send metrics and annotations from a training,
evaluation, or benchmark job to [Syvain Metrics](https://metrics.syvain.com/).
Use `syvain-metrics-api-client` or `syvain-metrics` cli (available in npm) to read stored metrics.

## Install

```bash
uv add syvain-metrics-collector
```

## Collect an experiment

```python
from syvain_metrics_collector import Collector

collector = Collector(api_key="ak_org_...")
experiment = collector.experiment(
	slug="mamba-run-001",
	description="Baseline mamba training run",
	folder_id="00000000-0000-0000-0000-000000000000",
	meta={
		"model": "mamba",
		"dataset": "internal-v1",
		"seed": 7,
		"config": {"batch_size": 32, "learning_rate": 0.0003},
	},
)

with experiment.run():
	for step in range(1, 1_001):
	  # run actual training
		loss = 1.0 / step

		if step == 1 or step % 10 == 0:
			experiment.metric(
				"loss",
				loss,
				step=step,
				metadata={"split": "train"},
			)

	experiment.annotation(
		"Checkpoint saved",
		metadata={"path": "checkpoints/mamba-run-001/step-999.pt"},
	)

experiment.flush_or_raise()
```

`metric()` and `annotation()` enqueue data. The collector sends queued batches
in the background. `experiment.run()` records the lifecycle and attempts a
best-effort flush when the block exits. The final `flush_or_raise()` fails the
job if queued data cannot be delivered or the collector previously evicted
events after exceeding `max_queue_items`. Capacity drop counts persist for the
collector's lifetime, including after a successful queue drain.

Use exactly one `flush_or_raise()` after the `run()` block. Do not call
`flush()` or `flush_or_raise()` from the training loop, evaluation loop,
reporting branch, or checkpoint branch.

## Collect only measurements the experiment needs

Define the evidence before adding metrics. Each metric must be required to
answer the experiment's question or to interpret training health. Do not emit
every intermediate, tensor statistic, layer value, or runtime diagnostic.

Report training metrics at a planned cadence. Aggregate device tensors first,
then convert them to Python numbers only when reporting. This avoids a device
synchronization on every microbatch.

## Put values in the right field

| Value | Field |
| --- | --- |
| Stable run identity and configuration | Experiment `meta` |
| Numeric measurement | Metric `value` |
| Training or evaluation progress | Metric `step` |
| Bounded category used to group a series | Metric `metadata` |
| Unique event details, paths, hashes, IDs, and text | Annotation `metadata` |

Use one stable metric name for one quantity and unit. Keep the same name across
splits, datasets, stages, devices, and ranks:

```python
experiment.metric("loss", train_loss, step=step, metadata={"split": "train"})
experiment.metric("loss", valid_loss, step=step, metadata={"split": "valid"})
```

Do not encode dimensions in the name:

```python
# Wrong
experiment.metric(f"{stage}/{split}/loss", loss, step=step)

# Correct
experiment.metric(
	"loss",
	loss,
	step=step,
	metadata={"stage": stage, "split": split},
)
```

## Keep metric metadata low-cardinality

Metric metadata is a flat `str -> str` mapping. Use it only for bounded
categories needed to compare series, such as `split`, `dataset`,
`training_stage`, `device`, `rank`, or optimizer parameter group.

Every distinct metadata mapping creates a separate series. The product of all
dimension values, including missing-key variants, must stay at or below 4,096
series per metric in one experiment. For example, 8 stages, 3 splits, and 16
ranks produce 384 series.

Never put steps, epochs, timestamps, paths, sample or request IDs, hashes, free
text, numeric measurements, or serialized objects in metric metadata. Put
progress in `step`, stable configuration in experiment `meta`, and unique
details in annotations.

The client accepts at most 32 metadata keys, 128 UTF-8 bytes per key, 512 UTF-8
bytes per value, and 4,096 UTF-8 bytes in the canonical JSON mapping. It
validates these limits before enqueueing the metric. Metric values must be
finite numbers; the client logs and drops non-finite values.

## Use test collectors

Use `NoopCollector` when a test only needs the collector interface. Use
`JsonlCollector(path=...)` when a local run needs inspectable JSONL output. Both
collectors use the same experiment, metric, annotation, and lifecycle calls as
`Collector`.
