Metadata-Version: 2.1
Name: gpumanager
Version: 0.3.4
Summary: Lightweight NVIDIA GPU utilization sampler, Slack reporter, and systemd timer helper
Author: OpenAI Codex
License: MIT
Project-URL: Website, https://happilee12.github.io/gpu-util-webhook/
Keywords: gpu,nvidia,slack,systemd,monitoring,cli
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: System Administrators
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: System :: Monitoring
Classifier: Topic :: System :: Systems Administration
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: tomli >=1.1.0 ; python_version < "3.11"
Requires-Dist: backports.zoneinfo >=0.2.1 ; python_version < "3.9"

# gpumanager

`gpumanager` is a lightweight Python CLI that samples NVIDIA GPU utilization, stores one CSV snapshot per sample, averages utilization over a reporting window, and posts GPU-wise summaries to Slack through an incoming webhook.

One sampler feeds any number of report schedules. A realtime ping every 10 minutes, a daily average, and a quarterly summary can all run side by side from the same collected data.

- Website: https://happilee12.github.io/gpu-util-webhook/
- PyPI: https://pypi.org/project/gpumanager/

## Features

- Samples NVIDIA GPU utilization with `nvidia-smi`
- Stores one CSV file per sample, aggregated by GPU UUID
- Any number of report schedules, each with its own cron time and aggregation window
- Rolling windows (`last 7d`) or calendar-anchored windows (`since the start of this quarter`)
- Sends reports to Slack via incoming webhook
- Installs and reconciles system-wide `systemd` services and timers
- Interactive configuration, minimal dependencies, close to the standard library

## Requirements

- Linux with `systemd`
- Python 3.8 or newer
- NVIDIA GPU with `nvidia-smi` in `PATH`
- A Slack incoming webhook URL — see the [Slack documentation](https://docs.slack.dev/messaging/sending-messages-using-incoming-webhooks/)

Python 3.8 and 3.9 pull in small compatibility dependencies automatically (`tomli`, `backports.zoneinfo`).

## Installation

`pipx` is recommended: it keeps `gpumanager` in its own virtualenv while exposing the command globally.

```bash
pipx install gpumanager
pipx ensurepath        # run once if the command is not found
source ~/.bashrc
```

With plain `pip`:

```bash
pip install --user gpumanager
```

From a local checkout, to test a build before publishing it:

```bash
python3 -m build
pipx install --force dist/gpumanager-<version>-py3-none-any.whl
```

Verify which build is actually on `PATH`:

```bash
gpumanager --help
python3 -c "import gpumanager; print(gpumanager.__version__)"
```

If a project virtualenv is active, its own copy shadows the `pipx` one. Run `deactivate` first, or call `~/.local/bin/gpumanager` directly.

## Quick Start

```bash
gpumanager init                      # answer the prompts, add one or more reports
gpumanager test-sample               # collect one sample now
gpumanager test-report               # send it to Slack now
gpumanager install-systemd --enable-now
gpumanager status
```

`init` shows the current server time and cron examples while asking for each report's schedule. If timers are already installed, it rewrites and reloads them so changes take effect immediately.

On a multi-user machine, pass the account the services should run as:

```bash
gpumanager install-systemd --enable-now --run-user "$USER"
```

## How It Works

```
gpumanager-sample.timer  ──▶  nvidia-smi  ──▶  one CSV per sample in csv_dir
                                                      │
                          ┌───────────────────────────┴───────────────────────────┐
                          ▼                                                       ▼
        gpumanager-report-daily.timer                        gpumanager-report-quarterly.timer
        averages the last 1d  ──▶ Slack                      averages since Jul 1 ──▶ Slack
```

Sampling and reporting are separate timers. Reports only read the CSV files, so adding a report costs no extra sampling and no extra disk space. Disabling sampling leaves every report with nothing to aggregate.

## Configuration

Configuration is searched in this order:

1. Path passed with `--config`
2. `GPUMANAGER_CONFIG`
3. `~/.config/gpumanager/config.toml`
4. `/etc/gpumanager/config.toml`

```toml
[slack]
webhook_url = "https://hooks.slack.com/services/..."

[storage]
csv_dir = "/var/lib/gpumanager"

[sample]
interval = "1m"

[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"

[[report]]
name = "quarterly"
report_time = "0 9 1 1,4,7,10 *"
interval = "since:quarter"

[general]
timezone = "Asia/Seoul"
server_name = "AICA_H100"
```

| Key | Meaning |
| --- | --- |
| `slack.webhook_url` | Slack incoming webhook, shared by every report |
| `storage.csv_dir` | Where samples are written; must be writable by the service user |
| `sample.interval` | How often a sample is taken: `7s`, `30s`, `2m`, `15m`, `1h` |
| `report.name` | Unique schedule name, also the systemd unit name |
| `report.report_time` | When the report is sent, as a 5-field cron string |
| `report.interval` | How far back the average reaches |
| `general.timezone` | IANA name; report times and `since:` boundaries resolve against it |
| `general.server_name` | Shown in the Slack message header |

### Report schedules

Each `[[report]]` block (note the double brackets) is one schedule, and any number of them can be configured.

`name` is required and must be unique. It may contain lowercase letters, digits, `-` and `_` only, up to 32 characters, because it becomes part of a systemd unit name installed with `sudo` (`gpumanager-report-<name>.timer`).

A single `[report]` table from an older version is still read correctly, as one report named `default`.

### `report_time` — when to send

A 5-field cron string: minute, hour, day-of-month, month, day-of-week.

| Schedule | Cron |
| --- | --- |
| Every day at 09:00 | `0 9 * * *` |
| Every hour | `0 * * * *` |
| Every 10 minutes | `*/10 * * * *` |
| Every Monday at 09:00 | `0 9 * * 1` |
| First day of each quarter at 09:00 | `0 9 1 1,4,7,10 *` |

### `interval` — how far back to average

A rolling duration, counted back from the moment the report is sent:

| Value | Window |
| --- | --- |
| `30m`, `1h`, `12h` | The last 30 minutes / 1 hour / 12 hours |
| `1d`, `7d` | The last 24 hours / 7 days |

Or a calendar anchor with the `since:` prefix, starting at 00:00 on the boundary day:

| Value | Window |
| --- | --- |
| `since:day` | Since midnight today |
| `since:week` | Since Monday of this week |
| `since:month` | Since the 1st of this month |
| `since:quarter` | Since the most recent Jan 1 / Apr 1 / Jul 1 / Oct 1 |
| `since:year` | Since January 1st |
| `since:2026-01-15` | Since a fixed date |

The difference matters at boundaries: on September 21st, `7d` reaches back into the previous quarter, while `since:quarter` stops cleanly at July 1st. The resolved start is printed in the Slack message, so the message itself says what was counted.

## Managing Report Schedules

Add a schedule. The new timer is installed and started; existing schedules are untouched.

```bash
gpumanager add-report realtime --report-time "*/10 * * * *" --interval 1h
gpumanager add-report                   # prompts for name, cron and window
```

Remove a schedule. The block is dropped from the config, the timer is disabled and its unit files are deleted, so the report also disappears from `gpumanager status`.

```bash
gpumanager remove-report realtime
gpumanager remove-report realtime --yes # skip the confirmation
```

The last remaining report cannot be removed — use `uninstall-systemd` to remove everything instead.

Stop a report without deleting it. The `[[report]]` block stays, and `reload` will not bring the timer back.

```bash
gpumanager disable-report --report weekly
sudo systemctl enable --now gpumanager-report-weekly.timer   # to resume
```

Editing the config by hand works too — the config file is the single source of truth:

```bash
$EDITOR ~/.config/gpumanager/config.toml
gpumanager reload
```

`reload` reconciles everything: new reports get their timers installed and started, deleted reports get their timers disabled and removed.

Send a report immediately:

```bash
gpumanager test-report                  # every report; asks first when several exist
gpumanager test-report --report daily   # just one
gpumanager test-report --yes            # every report, no question (for scripts)
```

Commands that act on every report at once ask for confirmation only when more than one report is configured and the terminal is interactive. The installed timers always target a single report by name, so they never wait for an answer.

## Automatic Scheduling

`gpumanager` does not collect anything in the background on its own. Install the system timers:

```bash
gpumanager install-systemd --enable-now
```

This writes to `/etc/systemd/system/` (via `sudo`):

- `gpumanager-sample.service` and `gpumanager-sample.timer`
- `gpumanager-report-<name>.service` and `gpumanager-report-<name>.timer`, one pair per `[[report]]` block

`add-report`, `remove-report`, `install-systemd` and `reload` all reconcile the installed units against the config file:

- timers for newly added reports are enabled and started
- units for reports no longer in the config are disabled and deleted
- the single unnamed `gpumanager-report.{service,timer}` pair from versions before 0.3.0 is replaced by `gpumanager-report-default.*`

Only newly added timers are enabled, so a schedule stopped with `disable-report` is not silently switched back on.

Stop sampling, which stops the data every report depends on:

```bash
gpumanager disable-sample
sudo systemctl enable --now gpumanager-sample.timer   # to resume
```

Remove every unit:

```bash
gpumanager uninstall-systemd
```

## Checking Status

```bash
gpumanager status
```

```json
{
  "config_path": "/home/master/.config/gpumanager/config.toml",
  "csv_dir": "/var/lib/gpumanager",
  "csv_dir_exists": true,
  "sample.interval": "1m",
  "reports": [
    {
      "name": "daily",
      "report_time": "0 9 * * *",
      "on_calendar": "*-*-* 09:00:00 Asia/Seoul",
      "interval": "1d",
      "timer": "gpumanager-report-daily.timer",
      "timer_installed": true,
      "next_trigger": "Tue 2026-03-24 09:00:00 KST; 22h left"
    }
  ],
  "sample_timer_installed": true,
  "sample.next_trigger": "Tue 2026-03-24 14:41:35 KST; 9s left"
}
```

`next_trigger` values are read by parsing the `Trigger:` line of `systemctl status <timer unit>`.

## Sampling and Report Format

Each sample writes one CSV named after its timestamp:

```text
2026-03-22T16-21-00.csv
```

```csv
timestamp,gpu_index,gpu_uuid,gpu_name,util_gpu
2026-03-22T16:21:00+09:00,0,GPU-aaa,NVIDIA A100,35
2026-03-22T16:21:00+09:00,1,GPU-bbb,NVIDIA A100,2
```

Reports average by GPU UUID over the window and round to two decimal places. Missing samples are ignored. The header is `[server_name/report_name]`; a report named `default` shows only the server name.

```text
[AICA_H100/quarterly] 2026.10.01 09:00:00 KST
Window: since 2026-07-01 00:00
GPU 0: 31.38%
GPU 1: 29.39%
GPU 2: 31.57%
GPU 3: 56.36%
```

Old CSV files can be deleted by range:

```bash
gpumanager delete-csv --start "2026-01-01 00:00:00" --end "2026-03-31 23:59:59"
```

## Command Reference

| Command | Purpose |
| --- | --- |
| `gpumanager init` | Interactive setup; walks through every setting and report |
| `gpumanager add-report [NAME] [--report-time CRON] [--interval WINDOW]` | Add one report schedule and install its timer |
| `gpumanager remove-report NAME [--yes]` | Delete one report schedule and its timer |
| `gpumanager test-sample` | Collect and store one sample now |
| `gpumanager test-report [--report NAME] [--yes]` | Aggregate and send to Slack now |
| `gpumanager status` | Configuration, installed timers and next run times, as JSON |
| `gpumanager install-systemd [--enable-now] [--run-user USER]` | Install and reconcile the system units |
| `gpumanager reload` | Re-apply the config to the installed units |
| `gpumanager disable-sample` | Stop the sampling timer |
| `gpumanager disable-report [--report NAME] [--yes]` | Stop report timers, keeping their config |
| `gpumanager uninstall-systemd` | Disable and delete every unit |
| `gpumanager delete-csv [--start S] [--end E] [--yes]` | Delete stored CSV files in a datetime range |

Every command accepts `--config PATH` to target a specific configuration file.

## Recommended Setups

Realtime activity ping:

```toml
[[report]]
name = "realtime"
report_time = "*/10 * * * *"
interval = "1h"
```

Daily average at 09:00:

```toml
[[report]]
name = "daily"
report_time = "0 9 * * *"
interval = "1d"
```

Weekly summary every Monday:

```toml
[[report]]
name = "weekly"
report_time = "0 9 * * 1"
interval = "7d"
```

Quarterly summary that starts exactly at the quarter boundary:

```toml
[[report]]
name = "quarterly"
report_time = "0 9 1 1,4,7,10 *"
interval = "since:quarter"
```

## Troubleshooting

**The test report arrives, but nothing comes at the scheduled time.**
The timers are probably not installed. Run `gpumanager status` and check `sample_timer_installed` and each report's `timer_installed`. Install them with `gpumanager install-systemd --enable-now`.

**Reports say `No GPU samples found in the selected window.`**
Either sampling is not running (`gpumanager status` → `sample.next_trigger`), or the service user cannot write to `csv_dir`, or the window is shorter than the sampling interval.

**A schedule was deleted from the config but still fires.**
Run `gpumanager reload`, which removes units for reports that are no longer configured.

## Upgrading

```bash
pipx upgrade gpumanager          # or: pip install --user --upgrade gpumanager
gpumanager reload
gpumanager status
```

Existing configuration files keep working across upgrades. A pre-0.3.0 single `[report]` table is read as one report named `default`, and its unit pair is replaced by `gpumanager-report-default.*` the next time units are installed. The config file itself is only rewritten when a command that changes it runs (`init`, `add-report`, `remove-report`).

## Version History

### 0.3.4

- Rewritten README: installation, configuration, schedule management, systemd behaviour, command reference and troubleshooting, plus this version history.

### 0.3.3

- `remove-report NAME` deletes one report from the config and removes its timer, so it also disappears from `status`. The last remaining report is protected.
- `interval` accepts calendar-anchored windows: `since:day`, `since:week`, `since:month`, `since:quarter`, `since:year` and `since:YYYY-MM-DD`, in addition to rolling durations such as `7d`. Useful for quarterly or month-to-date reports that must start on a boundary rather than N days ago.
- The Slack `Window:` line shows the resolved start (`since 2026-07-01 00:00`) instead of only the raw value.
- Window validation happens in one place, so an invalid `interval` is rejected when the config is read rather than when the report is sent.

### 0.3.2

- `add-report [NAME] [--report-time CRON] [--interval WINDOW]` adds a schedule without rerunning `init`, and installs and starts its timer.
- `test-report` and `disable-report` ask for confirmation before acting on every report at once, with `--yes` to skip. Non-interactive runs, including the systemd services, are never blocked by the prompt.
- Only newly added timers are enabled during a reconcile, so `disable-report` is not undone by a later `reload`.
- If the systemd step fails after the config was written, the error explains that `gpumanager reload` finishes the job.

### 0.3.0

- Multiple report schedules. `[[report]]` blocks replace the single `[report]` table; each has a `name`, `report_time` and `interval`.
- One systemd unit pair per report (`gpumanager-report-<name>.{service,timer}`), reconciled against the config by `install-systemd` and `reload`: new timers installed, obsolete ones disabled and deleted.
- `test-report --report NAME` and `disable-report --report NAME` act on a single schedule.
- Slack messages include the report name in the header: `[AICA_H100/weekly]`.
- `status` reports a `reports` array with per-report `on_calendar`, `timer_installed` and `next_trigger`, replacing the flat `report.*` keys.
- Configs from earlier versions load unchanged as a single report named `default`, and the old unnamed unit pair is replaced automatically.

### 0.2.x and earlier

- Single report schedule, `nvidia-smi` sampling into per-sample CSV files, Slack webhook delivery, systemd sample and report timers, `--run-user` for the service account.

## Notes

- `sample.interval` controls collection; `report.interval` controls only how far back each report averages
- `since:` boundaries and cron times both resolve against `general.timezone`; weeks start on Monday
- The README is used as the package long description, so this guide also appears on the PyPI project page
