Metadata-Version: 2.5
Name: trainmeter
Version: 0.0.3
Summary: Where your training FLOPs go and how well the GPUs are used: loss, FLOPs, MFU and hardware counters for a wrapped training run.
Project-URL: Homepage, https://github.com/Almaz-KG/trainmeter
Project-URL: Repository, https://github.com/Almaz-KG/trainmeter
Project-URL: Issues, https://github.com/Almaz-KG/trainmeter/issues
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: flops,gpu,mfu,nvml,pytorch,training
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: System :: Monitoring
Requires-Python: >=3.10
Provides-Extra: gpu
Requires-Dist: nvidia-ml-py>=12.535; extra == 'gpu'
Description-Content-Type: text/markdown

# trainmeter

**Where your training FLOPs go, and how well your GPUs are used.**

Run `tm train.py` instead of `python train.py`.
trainmeter starts your command, watches the node from outside, and serves a live dashboard on localhost.
The dashboard shows what a training run needs to be judged: training and validation loss, FLOPs spent and required, MFU with its FLOPs convention spelled out, and how busy the GPUs really are, from `GPU-Util` down to MFU, next to arithmetic intensity.
Your training code stays as it is.

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="docs/images/dashboard-dark.png">
  <img alt="The trainmeter dashboard: progress, throughput, MFU and OFU, goodput and the loss curve of a run on four H100s" src="docs/images/dashboard-light.png">
</picture>

<sub>The dashboard on a synthetic run.</sub>

Status: early scaffold.
`tm` runs any command untouched, records the run, and turns what a PyTorch loop does into tokens, FLOPs spent and required, ETA, MFU (given a device) and the losses you `emit()`.
GPU counters (NVML and GPM: the utilization ladder, OFU, observed FLOPs, arithmetic intensity) and `tm doctor` are built and tested against fakes, not yet on a real GPU.
Series your loop already logs to wandb, TensorBoard or MLflow are picked up without a code change (`--no-tap` turns it off), and DataLoader waits, checkpoint saves and eval spans feed a "where the time goes" view.
One real step is also counted with `FlopCounterMode` (shown beside MFU as the `counted` convention), and layers and width are read from a model config when there is one.
Each run ends with a run passport (`passport.md` and `passport.json`) for a model card.
The dashboard, one page of progress, tiles and charts with the run details and model card folded away, is served on localhost while the run goes; it has been checked in a browser on synthetic data, not on a real run.
Not yet used on a real run.
To try it, `examples/` has a tiny GPT to run under `tm` on a laptop and a run planner that needs no GPU.

## Install

trainmeter is on [PyPI](https://pypi.org/project/trainmeter/):

```sh
# into the training environment, next to torch: this is what lets tm see inside the program
uv add "trainmeter[gpu]"
# or as a standalone tool, enough for tm view, tm doctor and tm passport
uv tool install trainmeter
```

The `gpu` extra pulls in `nvidia-ml-py` for the GPU counters and can be left out on a machine with no NVIDIA GPU.
Without `uv`, `pip install "trainmeter[gpu]"` does the same.

## How it will be used

```sh
tm train.py                                 # instead of: python train.py
tm --budget-tokens 20B train.py             # declare the plan: FLOPs required, progress, ETA
tm -- torchrun --nproc_per_node=8 train.py  # any command, one grammar
tm doctor                                   # what tm can see on this node
tm passport                                 # model-card passport of the latest run
tm export wandb --project p                 # upload a finished run to W&B
tm view                                     # replay the latest finished run in the dashboard
tm attach 12345                             # watch a program that is already running, from outside
```

`tm` prints one line with the dashboard URL and, at the end, one summary line.
The dashboard binds to loopback on the node, so a remote node needs an SSH tunnel (`ssh -L 8765:localhost:8765 node`).
`--port N` picks the first port to try and `--no-web` turns the dashboard off.

## How it works

- **From outside.**
  `tm` starts your command as a child process, reads the GPU counters (tensor, DRAM and SM activity via NVML GPM on Hopper and newer), and keeps the run log.
- **From inside, passively.**
  A small agent that `tm` injects into the training interpreter sees steps, batch shapes, parameter counts and data waits.
  It never synchronizes a device and never raises into your loop.
- **From your loggers.**
  The loss and any other series your loop already sends to wandb, TensorBoard or MLflow are picked up as they are.
  A loop with no tracker adds one `emit(train_loss=...)` line.
- **Physics, not prices.**
  The run log holds time, tokens, FLOPs, counters and the series your loop reports.
  Money is an optional overlay computed at view time, and a price is never an input.

`docs/architecture.md` has the design, and `docs/spec.md` defines every metric.

## Why the convention matters

On a 1.04B model with d 2048, 18 layers and context 1024, FLOPs per token are 5.83 G by the Hugging Face count, 6.06 G by Megatron, 6.28 G by TorchTitan and nanochat, and 6.69 G when embeddings are counted as matmuls.
Same run, MFU differing by 15%. trainmeter always says which one it used.

## Develop

```sh
uv sync
uv run pytest
uv run ruff check .
uv run ruff format --check .
```

See `docs/spec.md` for what trainmeter shows and records, `docs/architecture.md` for how it is built, and `CLAUDE.md` for the invariants.

## License

Apache-2.0.
