Metadata-Version: 2.4
Name: landauer-gap
Version: 0.4.0
Summary: Like `time`, but for energy: measure the GPU, CPU and DRAM energy, cost and carbon of any command or local model.
Author: Bharat Sharma
License-Expression: MIT
Project-URL: Homepage, https://landauer-gap.vercel.app
Project-URL: Leaderboard, https://landauer-gap.vercel.app/#/board
Project-URL: Source, https://github.com/bsharma173860-oss/d3-research/tree/main/apps/landauer-gap/joules
Project-URL: Issues, https://github.com/bsharma173860-oss/d3-research/issues
Keywords: energy,gpu,nvml,rapl,carbon,llm,benchmark,landauer
Classifier: Development Status :: 4 - Beta
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Programming Language :: Python :: 3
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Environment :: Console
Classifier: Topic :: System :: Monitoring
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# joules

**Like `time`, but for energy.** Wrap any command and get the real GPU, CPU and DRAM energy it
used, what that cost, its carbon, and how far it sits above the Landauer limit of physics.
Benchmark a local model and get joules per token, and how that compares with an API's price.

Illustrative output:

```
$ joules -- python train.py
joules · python train.py · 2.31 h · exit 0
  GPU 0 NVIDIA H100 80GB HBM3   4.91 MJ    590 W avg  nvml counter
  CPU package 0                  612 kJ     74 W avg  rapl counter
  DRAM (package 0)              98.7 kJ     12 W avg  rapl counter
  ──────────────────────────────────────────────────────────
  total  5.62 MJ = 1.56 kWh
  cost 0.156 at 0.1/kWh · CO₂ 577 g at 370 g/kWh
```

```
$ joules bench ollama:llama3.1:8b --api-price 0.20
  per token      3.74 J   ·   1.04 kWh per 1M tokens
  electricity    0.104 per 1M output tokens at 0.1/kWh
  vs API         0.2 per 1M → electricity is 1.9× cheaper than the API (hardware cost not included)
```

No dependencies: only the Python standard library and your GPU driver.

## Install

```bash
pipx install landauer-gap                    # or: pip install landauer-gap
pipx install "git+https://github.com/bsharma173860-oss/d3-research#subdirectory=apps/landauer-gap/joules"
joules devices                               # what can be measured on this machine
```

## Commands

| Command | What it does |
|---|---|
| `joules -- CMD …` / `joules run -- CMD …` | Measure a command. Its exit code passes through, so it works in CI. |
| `joules bench ollama:MODEL` | Energy per token of a model served by Ollama. |
| `joules bench openai:MODEL --url http://host:8000/v1` | Same for any OpenAI-compatible server: vLLM, llama.cpp, LM Studio, TGI. |
| `joules run --track PROJECT -- CMD …` | Measure and save the run to your private [Tracking page](https://landauer-gap.vercel.app/#/track). Also on `bench`. |
| `joules bench … --share` | Add the result to the public [leaderboard](https://landauer-gap.vercel.app/#/board). |
| `joules share FILE` | Share a result saved earlier with `--out FILE`. |
| `joules ci --budget 10 -- CMD …` | Energy check for a pull request: this change vs the base branch. |
| `joules proxy --upstream URL` | Energy receipt on every response of a model server. |
| `joules devices` | List the meters joules found, and why any are missing. |

Useful options: `--json` / `--out FILE` (machine-readable result), `--subtract-idle 5` (measure the
idle machine first and also report energy above idle), `--gpus 0,1`, `--price`, `--grid fr` or
`--ci 55`, `--pue 1.2`.

## Tracking your runs over time

```bash
export LANDAUER_API_KEY=lgk_…
joules run --track llama-ft --params 8 --tokens 2 -- python train.py
joules --track nightly -- ./train.sh            # short form
```

Each tracked run appears on the Tracking page within seconds: Landauer gap over time, energy per run
and totals for the project. joules sends the energy, time, cost, CO₂, hardware, operations and gap,
and a short label: the program and script name (`python train.py`) or your `--label`. It never sends
arguments, inline code, outputs or file contents. A tracking problem is reported but never changes
the command's exit code, so it is safe in CI.

## Energy check on every pull request

`joules ci` runs a command on your checkout and on the base branch (a temporary git worktree),
alternating runs so drift hits both sides, and compares the medians. It writes a Markdown summary,
can post one pull request comment that it updates on each push, and fails when energy grows by more
than `--budget` percent. A change smaller than the run-to-run spread is reported as noise, never as a
failure.

```bash
joules ci --base main --runs 3 --budget 10 -- pytest -q
```

In GitHub Actions use the [PR energy check action](../integrations/pr-energy-check/README.md). On
hosted runners, which have no energy counters, energy is estimated from CPU time and the comment says
so; on a self-hosted runner with a GPU or RAPL it is measured. `joules run --estimate` uses the same
estimate when a machine has no counters.

## Energy receipts for AI responses

`joules proxy` sits in front of a model server and measures the hardware while each request runs.
Point your client at the proxy instead of the server; nothing else changes.

```bash
joules proxy --upstream http://localhost:11434          # Ollama
joules proxy --upstream http://localhost:8000 --subtract-idle 5   # vLLM, measured above idle
```

```
$ curl -si localhost:8787/v1/chat/completions -d '{"model":"llama3.1:8b","messages":[…]}'
X-Energy-Joules: 41.2
X-Energy-Output-Tokens: 212
X-Energy-Joules-Per-Token: 0.194
X-Energy-CO2-Grams: 0.00423
X-Energy-Receipt: /joules/receipts/r-000017
```

- Streaming responses (SSE or Ollama's NDJSON) pass through untouched and carry `X-Energy-Receipt`;
  the receipt at that address is complete when the stream ends.
- Tokens come from the response (`usage`, or Ollama's `eval_count`). Without them, streamed pieces
  are counted and the receipt says `tokens_exact: false`.
- Meters measure whole devices, so in each sampling interval the energy is shared equally between
  the requests running at that moment. With `--subtract-idle` receipts also carry energy above idle.
- `/joules/receipts` lists the newest receipts, `/joules/metrics` serves Prometheus metrics, and
  `/joules/health` says what is measured.
- `--track PROJECT` sends one summary per model every 5 minutes to your Tracking page: never
  prompts or outputs.
- It listens on 127.0.0.1 by default. `--host 0.0.0.0` exposes it, and your model server, to your
  network.

## The leaderboard

```bash
export LANDAUER_API_KEY=lgk_…                # free key: landauer-gap.vercel.app → API
joules bench ollama:llama3.1:8b --subtract-idle 5 --share
```

The leaderboard ranks models and machines by measured joules per output token, using the median of
everyone's runs. joules sends only the model name, hardware name and count, energy per token,
tokens per second, average power, output tokens, whether idle power was subtracted, which meters
were used and the joules version. It never sends prompts, outputs, host names or paths. Simulated
runs are refused, and you can remove your own entries on the Leaderboard page.

## Estimate vs measured, and calibration

Tell joules the size of the training job and it compares the measurement with the Landauer Gap
estimate, then works out what your hardware really achieved:

```
$ joules run --params 1 --tokens 2 -- python finetune.py        # 1B model, 2B tokens, one H100
  estimate vs measured  17 MJ estimated (40% MFU, 80% draw) → 14.8 MJ measured (-13%)
  your hardware ran at 42.5% effective utilisation and 74% of rated power
  open in the calculator with these numbers: https://landauer-gap.vercel.app/?s=…
```

The link opens the website calculator with your measured utilisation and power draw, so every
future estimate starts from your real hardware instead of a textbook default. If the job you
describe could not fit in the measured time, joules says so instead of reporting nonsense.

## What it measures, and how

| Hardware | Source | Kind |
|---|---|---|
| NVIDIA GPUs | NVML (driver library, via ctypes) | energy counter on Volta and newer; power sampling on older cards |
| AMD GPUs (Instinct, Radeon) | amdgpu hwmon (`energy1_input`, else `power1_average`) | counter or sampling |
| Intel GPUs (Arc, Flex, Max) | i915 / xe hwmon `energy1_input` | energy counter |
| Intel and AMD CPUs, DRAM | RAPL (`/sys/class/powercap`) | counter, wraparound handled |
| Apple Silicon Macs | `powermetrics` (needs sudo) | sampling: CPU + GPU + Neural Engine |
| Intel Macs | `powermetrics` (needs sudo) | sampling: CPU package (cores, integrated GPU, DRAM) |

Meters cover whole devices, so anything else running on the same GPU or CPU is included. Use
`--subtract-idle`, and keep the machine otherwise quiet, for the cleanest numbers. Power supply
losses, fans, networking and cooling are not measured; `--pue` adds a facility overhead.

`CUDA_VISIBLE_DEVICES` is respected, with indices, GPU UUIDs (as Kubernetes and Slurm set them) or MIG
slices (MIG measures the whole GPU). On AMD, `HIP_VISIBLE_DEVICES` / `ROCR_VISIBLE_DEVICES` and `--gpus` pick cards. Reading RAPL on recent Linux kernels needs root, or
`sudo chmod a+r /sys/class/powercap/intel-rapl:*/energy_uj`.

## Tests

```bash
PYTHONPATH=src python3 -m unittest discover -s tests -v
```

The real sysfs layouts are rebuilt in a temporary directory (including counter wraparound), the
NVML binding runs against a compiled stand-in for `libnvidia-ml.so.1`, and model servers are
local HTTP stand-ins. `JOULES_SIMULATE="gpu:H100:650,cpu:package 0:80"` runs everything on
constant simulated power for demos; results are always labelled SIMULATED.
