Metadata-Version: 2.5
Name: bdbackup
Version: 0.6.3
Summary: File and MySQL/MariaDB backups with verification, retention, and guided recovery
Project-URL: Homepage, https://github.com/ayder/bdbackup
Project-URL: Repository, https://github.com/ayder/bdbackup
Project-URL: Issues, https://github.com/ayder/bdbackup/issues
Author-email: Sinan Alyuruk <sinan@dbsmedya.com>
License-Expression: MIT
License-File: LICENSE
Keywords: backup,mariabackup,mysql,tar,xtrabackup
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: System Administrators
Classifier: Operating System :: POSIX
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: System :: Archiving :: Backup
Requires-Python: >=3.12
Requires-Dist: click>=8.0
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest-cov; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Provides-Extra: mysql
Requires-Dist: pymysql>=1.0; extra == 'mysql'
Description-Content-Type: text/markdown

# bdbackup

**Backups by design.**

`bdbackup` is a Python library and command-line tool for file and MySQL/MariaDB
backups with verification, retention, and guided recovery. Define repeatable jobs,
track their outcomes, and prepare recovery copies with explicit safeguards at
each step.

Choose the backup method that fits your data:

- **file** backups from a plain-text template file into a tar archive
- **mysqldump** logical backups (gzip-compressed)
- **xtrabackup / mariabackup** physical full and incremental backups

Built for day-to-day system administration: TOML configuration, credential files,
verification before publication, process locking, atomic publication, and optional
SQLite history keep backup operations explicit and inspectable.

## Support and safety

Python 3.12+ on POSIX systems (Linux/macOS); Windows is not supported.
`tar` and `tar.gz` work on every supported Python; `tar.zst` requires Python 3.14
with Zstandard support. Database binaries are separate system dependencies.

Every built-in backend verifies staging output before publishing it. A missing
or unreadable file fails the backup and preserves the previous archive. A process
lock prevents concurrent writers to the same file destination or physical root.
`--no-verify` skips only the additional verification pass after publication.
Verification checks archive readability or physical checkpoints; it does not
replace testing a restore. File backups are not filesystem snapshots: quiesce
applications or back up snapshots when files can change during a run.

Safe relative symlinks, hardlinks, empty directories and ordinary directory
permissions/timestamps round-trip. Restore rejects escaping paths/links and
special files on all supported Python versions. Ownership and privileged
permission bits are intentionally not restored. Use a destination that is not
being modified by another process during extraction.

## Installation

Install the current code from the original GitHub repository:

```bash
pip install "git+https://github.com/ayder/bdbackup.git"
```

For versions published to PyPI:

```bash
pip install bdbackup
```

With optional MySQL metadata support:

```bash
pip install "bdbackup[mysql]"
```

For development:

```bash
git clone https://github.com/ayder/bdbackup.git
cd bdbackup
pip install -e ".[dev]"
```

## Global options

Place the global option before the command: `bdbackup --logging DEBUG file ...`.
Levels: `DEBUG`, `INFO`, `WARNING`, `ERROR`.
The CLI uses these exit codes:

| Code | Meaning |
|------|---------|
| 0 | Success |
| 1 | Backup/verify failure, or configuration helper checks failed/incomplete |
| 2 | Usage error |
| 3 | Lock held, or retention refused an unsafe/incomplete scan |

## File backup

Create a template file listing paths (one per line, `#` for comments,
whitespace allowed):

```text
# backup-template.txt
Documents
Videos
data/projects
```

Run the backup:

```bash
bdbackup file --template backup-template.txt -d /backup/files/daily
```

Gzip-compress:

```bash
bdbackup file --template backup-template.txt -d /backup/files/daily --format tar.gz
```

Options:

| Flag | Description |
|------|-------------|
| `--template` | Template file listing paths to archive |
| `-d, --dst` | Destination archive path (without extension) |
| `-c, --chdir` | Source path: resolve template paths and relative excludes against it |
| `-f, --format` | Archive format: `tar`, `tar.gz`, `tar.zst` (Python 3.14+) |
| `-x, --exclude` | Exact resolved path to exclude (repeatable) |
| `--exclude-pattern` | Glob pattern to exclude (repeatable) |
| `--exclude-template` | Named exclusion template, e.g. `python-dev` (repeatable, comma-separated allowed) |
| `--follow-symlinks` | Follow symbolic links when archiving |
| `--timestamp / --no-timestamp` | Write `<dst>-<UTC YYYY-MM-DD-HHMMSS><ext>` instead of replacing one archive; an existing name is never overwritten (default: off) |
| `--dry-run` | List what would be archived without writing anything |
| `--verify / --no-verify` | Verify the archive after creation (default: on) |
| `--logging` | Log level: `DEBUG`, `INFO`, `WARNING`, `ERROR` |

Without `--timestamp`, every run replaces the same archive, so only the newest
run stays restorable. With it, each run writes a new archive such as
`daily-2026-09-29-020000.tar.gz`, and a run that would reuse an existing name
fails instead. Config jobs set `timestamp = true`.

Restore a file archive:

```bash
bdbackup restore /backup/files/daily.tar -d /restore/here
```

### Exclusion templates

`--exclude-template` applies a named bundle of gitignore-style exclusion
patterns so you don't have to repeat common artifact rules:

```bash
bdbackup file --template backup-template.txt -d /backup/files/daily \
    --exclude-template python-dev
```

The built-in `python-dev` template skips `__pycache__/`, `*.pyc`, `.venv/`,
`venv/`, `uv.lock`, `Pipfile.lock`, `poetry.lock`, `*.egg-info/`, `build/`,
`dist/`, and common tool caches (`.mypy_cache/`, `.pytest_cache/`,
`.ruff_cache/`, `.tox/`, ...). Templates are repeatable and comma-separated
values work too; they combine with `-x/--exclude` and `--exclude-pattern`.

Patterns use gitignore-style syntax: `*.pyc` matches at any depth, a trailing
`/` matches directories only, patterns containing a `/` (e.g.
`tests/containers/*img`) match relative to the backup root, a leading `/`
anchors to the root, and `!` negates a previous pattern.

Custom templates are plain Python files in `~/.config/bdbackup/templates/`
(honours `$XDG_CONFIG_HOME` and `$BDBACKUP_CONFIG_DIR`) that self-register —
no existing code needs to change to add one:

```python
# ~/.config/bdbackup/templates/go_dev.py
from bdbackup.templates import ExclusionTemplate

TEMPLATE = ExclusionTemplate(
    name="go-dev",
    patterns=("vendor/", "vendor/**", "*.test", "go.work"),
    description="Go development artifacts",
)
```

## Database engines

`schedule` and `restore_root` are job metadata and are not passed to an engine backend.
Custom engines accepting arbitrary keyword arguments no longer receive `restore_root`.

Database engines are pluggable and live in per-database family packages
(`bdbackup/mysql/` today; `bdbackup/postgres/` is planned). Each engine maps a
config `type` name to a backend class; adding one never requires editing
existing wiring code:

| Engine | Config `type` | Backend | Notes |
|--------|---------------|---------|-------|
| `mysqldump` | `mysqldump` | `bdbackup.mysql.MySQLBackup` | Logical, gzip-compressed SQL dumps |
| `xtrabackup` | `xtrabackup` | `bdbackup.mysql.XtraBackup` | Percona XtraBackup / MariaDB mariabackup |

To add an engine, drop a self-registering module into either
`bdbackup/mysql/` (shipped with the package) or
`~/.config/bdbackup/engines/` (user-level, no package changes) — a versioned
xtrabackup variant or a mysql-shell `util.dump()` engine both follow the same
recipe:

```python
# ~/.config/bdbackup/engines/mysql_shell.py
from bdbackup.engines import EngineInfo

class MySQLShellDump:
    def __init__(self, out_dir: str = ".", **params): ...
    def backup(self, name=None): ...          # BackupBackend protocol
    def verify(self, result=None): ...
    def prune(self): ...

ENGINE = EngineInfo(
    name="mysql-shell",          # usable as type = "mysql-shell" in config
    backend=MySQLShellDump,
    description="MySQL Shell util.dump() backups",
    family="mysql",
)
```

A config `type` is validated against the live registry, so the new engine is
immediately usable from `config.toml`. MySQL-specific helpers shared by the
family (e.g. `mysql_cnf_file` for credential-safe defaults files) live in
`bdbackup/mysql/helpers.py`. A new database family means creating
`bdbackup/postgres/` and appending `"bdbackup.postgres"` to
`bdbackup.engines.FAMILIES` — the single deliberate modification point.

## mysqldump

Dump a single database:

```bash
bdbackup mysqldump --database mydatabase -o /backup/mysql/dumps -u root -p
```

Dump all databases:

```bash
bdbackup mysqldump --full -o /backup/mysql/dumps -u root -p
```

Parallel dumps of multiple databases:

```bash
bdbackup mysqldump --database db1,db2,db3 -o /backup/mysql/dumps -u root -p --jobs 4
```

Options:

| Flag | Description |
|------|-------------|
| `--database` | Database name(s), comma-separated (ignored when `--full`) |
| `-o, --out-dir` | Directory for the dump file |
| `-u, --user` | MySQL user |
| `-p, --password` | MySQL password; omit value to be prompted securely |
| `-h, --host` | MySQL host |
| `-P, --port` | MySQL port |
| `--options` | Comma-separated mysqldump options |
| `--full` | Add `--all-databases` and dump everything |
| `-j, --jobs` | Parallel dumps when multiple databases are specified |
| `--verify / --no-verify` | Verify the dump after creation (default: on) |

Use a bare `-p` to enter a password securely. Passing `-p PASSWORD` puts it in
bdbackup's own process arguments and potentially shell history. Child database
processes receive only a temporary credentials-file path; its contents are
quoted, its permissions are `0600`, and it is removed after the run.

Mysqldump retains `--single-transaction`, `--routines`, `--events` and `--triggers`
by default. `--options` adds options; explicit `--skip-*` flags can override
applicable defaults. `--full` retains these defaults and adds `--all-databases`.
Single-transaction consistency applies to transactional tables; quiesce writes
to nontransactional tables and avoid schema changes during a dump.

## xtrabackup

Full backup:

```bash
bdbackup xtrabackup full --database production -r /backup/mysql -u xtrabackup -p
```

Incremental backup (chains to the latest successful full):

```bash
bdbackup xtrabackup incremental --database production -r /backup/mysql -u xtrabackup -p
```

Prune old backups by retention days:

```bash
bdbackup xtrabackup prune --database production -r /backup/mysql --retention 7
```

Prepare a full backup or an incremental recovery point into a new directory:

```bash
bdbackup xtrabackup prepare /backup/mysql/production/2026-09-10/Full_ID \
    -r /backup/mysql/production -d /restore/production -u xtrabackup -p
```

Use the exact path printed by the backup command in place of `Full_ID`. Passing
an incremental path prepares its full and every prerequisite incremental up to
that point. Preparation copies sources to private working directories,
decompresses compressed copies, applies the increments in dependency order, and
publishes the recovery directory only on success. The destination must be new
and outside the backup root. Original backups remain available for new
incrementals and repeated recovery attempts.

Use XtraBackup matching your MySQL/Percona server series (8.0 with 8.0, 8.4 with
8.4); use `mariabackup` or `mariadb-backup` matching your MariaDB installation.
MariaDB preparation omits XtraBackup's `--apply-log-only` option. Compression is
**off by default**. For a compatible recent XtraBackup, select `--compress zstd`;
MariaDB's deprecated built-in compression accepts only `quicklz` and requires
`qpress` for decompression. Compatibility must be established with an actual
recovery test for the exact server and backup binary versions in use.

New physical backups record parent/full identities and LSNs in `bdbackup.json`.
Incrementals live under `DATE/Incremental/FULL_ID/UNIQUE_ID`, and cannot attach to
another full taken on the same day. Older backups without this metadata require
a new full before taking further incrementals; full backups can still be
prepared as recovery copies. Retention removes complete dated chains, preserves
the newest successful full's date, and refuses to prune without a valid full.

Options:

| Flag | Description |
|------|-------------|
| `mode` | `full`, `incremental`, `prune`, or `prepare` |
| `--database` | Database name (used for directory naming) |
| `-r, --root` | Backup root directory |
| `-u, --user` | MySQL user |
| `-p, --password` | MySQL password; omit value to be prompted securely |
| `-b, --binary` | `xtrabackup` or `mariabackup` |
| `--compress` | Compression algorithm (default: uncompressed) |
| `--compress-threads` | Compression threads (default: 4) |
| `--encrypt / --no-encrypt` | AES256 backup encryption (default: off; Percona only) |
| `--encrypt-key-file` | Required 32-byte key file when encrypting; also used by `prepare` |
| `--parallel` | Number of copy threads (default: 1) |
| `--throttle` | Limit I/O to this many IOPS |
| `--retention` | Days of backups to keep (default: 5); `0` disables engine deletion |
| `--verify / --no-verify` | Verify the backup after creation (default: on) |

A process lock prevents two backup runs from corrupting the same backup root.
Failed backups are written to a temporary directory first and cleaned up on
error, so a partial backup can never be mistaken for a complete one.

### Encrypted physical backups

Percona XtraBackup jobs can optionally encrypt full and incremental backups with
AES256. In the existing job's TOML section, add:

```toml
encrypt = true
encrypt_key_file = "/etc/mysql/xtrabackup.key"
```

TOML uses `true`/`false`, not `yes`/`no`. The key file must already exist, be
readable by the backup account, and contain exactly 32 bytes. To create a new
key once (this command refuses to overwrite an existing key):

```bash
sudo python3 - <<'PY'
import os
from pathlib import Path

key_path = Path("/etc/mysql/xtrabackup.key")
fd = os.open(key_path, os.O_WRONLY | os.O_CREAT | os.O_EXCL, 0o600)
with os.fdopen(fd, "wb") as key_file:
    key_file.write(os.urandom(32))
PY
```

Keep this key outside the backup tree and store a separate secure copy: losing
it makes its encrypted backups unrecoverable. Take a new full backup when changing
keys or encryption settings; incrementals must use their parent's settings and
key. Encryption is not supported by the mariabackup backend.

Run the configured job normally, or use the direct command:

```bash
bdbackup run --config /opt/dbs/config.toml mysql-prod
bdbackup --config /opt/dbs/config.toml xtrabackup full \
  --database production --root /backup/mysql \
  --encrypt --encrypt-key-file /etc/mysql/xtrabackup.key -u mysql -p
```

History records `encrypted = 1` in SQLite's `backup_runs` table for encrypted
attempts (`0` otherwise), including failed attempts. The `status` column indicates
whether an attempt succeeded. History also stores the algorithm and key-file
path for recovery, never the key itself. Existing history databases upgrade
automatically on the next backup; old entries remain unencrypted.

History restore uses the recorded key path. If the key has moved, provide its
new location:

```bash
bdbackup restore --config /opt/dbs/config.toml --backup-id 12 \
  --dst /restore/mysql-prod-12 --encrypt-key-file /secure/saved-xtrabackup.key --yes
```

Recovery decrypts each full/incremental work copy, then decompresses and prepares
it. Original backups remain encrypted; the recovery directory contains plaintext.
For manual `xtrabackup prepare`, supply `--encrypt-key-file` as well.
See [Percona's encryption documentation](https://docs.percona.com/percona-xtrabackup/8.4/encrypt-backups.html).

## Retention of flat files (`type = "retention"`)

This section handles flat backup files written by other tools. For bdbackup's
own backups, use `type = "gfs"` (see [Ledger and GFS](#ledger-and-gfs)), which
moves backups between storage tiers using the history ledger.

The `bdbackup retention` command applies a grandfather-father-son policy to
a backup tree: keep **every** backup for the last N days, then **one per ISO
week** for N weeks, then **one per calendar month** for N months. Tiers are
sequential and never overlap, so the retained set is deterministic and easy to
reason about.

```bash
# Dry run (default) -- lists what would be deleted, changes nothing
bdbackup retention --full-dir /backup/mysql/production \
    --incr-dir /backup/mysql/production/incr \
    --log-dir /backup/mysql/production/log \
    --daily 7 --weekly 4 --monthly 6 \
    --incr-days 7 --log-days 30 \
    --min-keep-fulls 2 --pick first

# Actually delete
bdbackup retention ... --apply
```

| Tier | Flag | Meaning |
|------|------|---------|
| Daily | `--daily N` | Keep every backup of the last N calendar days |
| Weekly | `--weekly N` | Then one backup per ISO week, for N weeks |
| Monthly | `--monthly N` | Then one backup per calendar month, for N months |
| Incrementals | `--incr-days N` | Keep incrementals for N days, never past their full |
| Logs | `--log-days N` | Keep log files for N days |
| Safety | `--min-keep-fulls N` | Never leave fewer than N newest fulls |
| Pick | `--pick first/last` | Which backup survives in a weekly/monthly bucket |

**Chain safety:** this flat-file policy expects one chronological backup
stream: each new full starts a chain, and subsequent increments continue that
chain until the next full. A retained incremental pins its full and every
preceding incremental in that chain, including prerequisites older than the age
window. A missing full makes an incremental an orphan. If a full/incremental
scan skips a file (for example while it is being written), expiration is deferred
because its dependencies are uncertain. Keep all files for a stream together;
arbitrary overlapping chains or missing intermediate backups require explicit
external metadata and are not supported by this filename-based policy.

This command handles flat backup files such as `full_*.mbi`, not the dated
XtraBackup directory tree. Use `bdbackup xtrabackup prune` for physical backups.

Only `--full-dir` is required. `--incr-dir` and `--log-dir` are optional.
`--full-glob`, `--incr-glob`, and `--log-glob` default to `full_*.mbi`,
`incr_*.mbi`, and `*.log`.

## Config-driven jobs

For scheduled or multi-job usage, define a TOML config file. Convention:
keep all bdbackup settings under `~/.config/bdbackup/` (honours
`$XDG_CONFIG_HOME` / `$BDBACKUP_CONFIG_DIR`) — `config.toml`, the
`files.template` path lists, `templates/*.py` exclusion presets, and
`engines/*.py` database engines.

```toml
# ~/.config/bdbackup/config.toml
[files-daily]
type = "file"
template_filename = "~/.config/bdbackup/files.template"
backup_dst = "/backup/files/daily"
chdir = "/srv/www"   # source path: template/exclude entries resolve against it
format = "tar.gz"  # tar.zst requires Python 3.14+
exclude_pattern = ["*.log", "node_modules"]
exclude_templates = ["python-dev"]

[mysql-prod]
type = "xtrabackup"
backup_root = "/backup/mysql/production"
user = "xtrabackup"
retention_days = 7
parallel = 2

[mysqldump-all]
type = "mysqldump"
out_dir = "/backup/mysql/dumps"
user = "backup"
options = ["--single-transaction", "--all-databases"]

[mysql-prod-retention]
type = "retention"
full_dir = "/backup/flat/full"
incr_dir = "/backup/flat/incr"
log_dir = "/backup/flat/log"
daily = 7
weekly = 4
monthly = 6
incr_days = 7
log_days = 30
min_keep_fulls = 2
pick = "first"
apply = false           # set true once the dry-run output looks right
```

Configuration paths expand `~`; relative config paths resolve against the
config file's directory. Template entries and `exclude` entries resolve against
`chdir` (or the process working directory if omitted). For a single-database
mysqldump job, set `database = "mydatabase"`; for all databases include
`--all-databases` in `options`.

Except for the optional `[history]` settings, each top-level table is one job;
its keys (except `type`, `schedule`, and `restore_root`) are passed
to the backend constructor, so they use Python-style underscores
(`exclude_templates`), not CLI dashes. See [config.toml.example](https://github.com/ayder/bdbackup/blob/main/config.toml.example)
for a fully annotated config with every option explained.

Run one job:

```bash
bdbackup run --config ~/.config/bdbackup/config.toml mysql-prod
```

Run every job:

```bash
bdbackup run --config ~/.config/bdbackup/config.toml --all
```

For an xtrabackup job, take an incremental using its own `backup_root`, credentials,
binary, and encryption settings:

```bash
bdbackup run --config ~/.config/bdbackup/config.toml mysql-prod --incremental
```

`--full` is the default; an incremental requires a successful full in that job's
root and chains to the latest successful incremental, if present. `--full` and
`--incremental` together exit 2. Selecting any non-xtrabackup job with
`--incremental` also exits 2 before any job runs. `--verify/--no-verify` applies to
both backup kinds. Validation and cron recommendations ignore the selector.

For example, schedule a weekly full on Sunday and incrementals on the other days:

```cron
0 2 * * 0 bdbackup --config /opt/dbs/config.toml run mysql-prod
0 2 * * 1-6 bdbackup --config /opt/dbs/config.toml run mysql-prod --incremental
```

Pruning after each full deletes whole dated chains older than `retention_days`.
Set `retention_days` longer than the interval between fulls to preserve the previous
full's incrementals. `retention_days = 0` disables engine deletion for the job,
both the automatic pruning after a full and `bdbackup xtrabackup prune`; use it
when another tool rotates the backups. History now records configured physical backups as
`xtrabackup-full` / `xtrabackup-incremental` (previously `xtrabackup`); update filters
that use the old type. Existing history rows are unchanged.

### Validate configuration and permissions

Check all configured jobs without creating a backup or running retention:

```bash
bdbackup --config /opt/dbs/config.toml --validate
# Or validate only one job:
bdbackup run --config /opt/dbs/config.toml mysql-prod --validate
```

Run validation as the OS account that will run your cron jobs. It checks backend
settings, required executables, destination access, SQLite/recovery directories,
file-template sources, and encryption key requirements. Missing directories are
reported with `mkdir -p` guidance and the required OS-user permissions. A missing
directory that the backend can create under a writable parent is reported as
creatable; validation does not create it.

For MySQL jobs, the installed `mysql`/`mariadb` client authenticates using the job's
credentials and reads `CURRENT_USER()`, server version, datadir and `SHOW GRANTS`.
Passwords are passed through a temporary mode-0600 options file, removed afterward.
No `CREATE USER` or `GRANT` is executed. Missing privileges produce SQL for an
administrator, for example:

```sql
GRANT RELOAD, BACKUP_ADMIN, REPLICATION CLIENT, PROCESS, LOCK TABLES
ON *.* TO 'xtrabackup_user'@'localhost';
GRANT SELECT ON `performance_schema`.`log_status`
TO 'xtrabackup_user'@'localhost';
```

The helper also checks the performance-schema tables used by current Percona
XtraBackup and prints `CREATE TABLESPACE` separately as an optional privilege for
importing individual tables. MariaDB Backup gets its own privilege recommendations.
See [Percona privileges](https://docs.percona.com/percona-xtrabackup/8.4/privileges.html)
and [MariaDB Backup privileges](https://mariadb.com/docs/server/server-usage/backup-and-restore/mariadb-backup/mariadb-backup-overview).

Checks cover direct grants; role-derived privileges and partial revokes may need
manual review. Physical datadir checks cover root-directory access, not every data
file or external tablespace. Custom mysqldump options, routine visibility, GTID
settings and exact server/binary version compatibility still need review.
An unreachable database or missing client is a failed/incomplete check, not a pass.
Retention validation checks settings and directory permissions without scanning
for deletions, even when `apply = true`. Helpers do not add SQLite history rows.

### Recommend cron entries

Inspect the invoking OS account's `crontab -l` and print suggested entries:

```bash
bdbackup --config /opt/dbs/config.toml --cron
bdbackup run --config /opt/dbs/config.toml mysql-prod --cron
# Both helpers can be used together:
bdbackup --config /opt/dbs/config.toml --validate --cron
```

Nothing is installed or edited. Defaults are daily backups starting at 02:00 and
retention starting at 04:00, staggered by 15 minutes within each group. Override
the time inside any job's existing TOML section:

```toml
[mysql-prod]
type = "xtrabackup"
backup_root = "/backup/mysql/production"
schedule = "30 1 * * *"

[mysql-prod-retention]
type = "retention"
full_dir = "/backup/flat/full"
schedule = "0 5 * * 0"
apply = false
```

`schedule` accepts five numeric cron fields with wildcards, lists, ranges and
steps. It is job metadata, never passed to the backup backend. The recommendations
use absolute executable/config paths, quote shell arguments, preserve the current
working directory and suggest the current PATH for cron. Times use the cron
daemon's timezone. Allow enough time for backups before retention; separate cron
entries do not establish a dependency.

Matching active `bdbackup run` entries for the same config/job (or `--all`) are
reported without suggesting duplicates. Commented entries and other configs do
not count. Shell wrappers such as `flock`, scripts, and system-wide cron files
are not inspected; review those separately. If crontab cannot be read, suggested
entries are still shown but the helper exits with status 1.

XtraBackup job entries run full backups, including their built-in pruning.
The helper recommends fulls only; add an incremental cron entry yourself, for example:

```cron
0 2 * * 1-6 bdbackup --config /opt/dbs/config.toml run mysql-prod --incremental
```

Adjust the recommended full entry to your intended full-backup schedule.

Separate retention jobs remain for flat backup files; `apply = false` stays a dry
run in cron too. Validation/cron helpers exit 0 when their checks pass, 1 for failed
or incomplete checks, and 2 for invalid command/configuration syntax.

## Backup history and guided restore

Enable SQLite history in your TOML configuration. No extra Python dependency
or database server is required:

```toml
[history]
database = "state/history.sqlite3"
restore_root = "/restore/bdbackup"
```

Both paths expand `~`; relative paths resolve against the configuration file.
`restore_root` defaults to `restores` beside that file. Omitting `[history]`
disables history. File jobs automatically exclude this live database and its
SQLite journal files; keep it outside backup inputs when possible. New history
databases are created with permissions `0600`.

`run --config` records every selected backup attempt. For direct commands,
place `--config` before the command:

```bash
bdbackup run --config config.toml files-daily
bdbackup --config config.toml mysqldump --database app,analytics --jobs 2
bdbackup history --config config.toml
bdbackup history --config config.toml --job files-daily --successful
```

Each record contains an ID, job name, backup type, UTC start/completion times,
status (`running`, `success`, or `failed`), and the successful artifact's absolute
path and size. Direct file commands use the destination name as the job name;
direct database commands use the database name. Parallel dumps get one record
per database, including when another dump fails. Physical full and incremental
commands record their respective types and the root/binary needed for preparation.
Credentials and raw error messages are not stored; failed rows contain only the
exception class. Retention, restore operations, and file dry runs are not backup
attempts and do not create rows. Direct Python backend calls are not automatically
recorded; applications can wrap them with `History.run`.

Every successful backup also records its unit and a SHA-256 checksum taken right
after the backup, which costs one extra read of the backup. A file's checksum is
the SHA-256 of its content. A directory's checksum is the SHA-256 of a manifest
with one line per regular file, `<sha256>  <relative path>`, sorted by path; a
symlink or special file inside a physical backup fails the backup. The unit is
what a rotation tool moves as one piece: the archive or dump file itself, or for
xtrabackup the dated directory that holds the full and all of its incrementals.
`bdbackup history` shows it as `unit <path>`, or `unit -` for records without one.
For units that GFS manages, the line continues with `| stage <stage path> |
locations <path>, …`, or `| deleted <time>` once GFS has removed the unit.
Reproduce a directory checksum on Linux (use `shasum -a 256` on macOS):

```bash
cd /backup/mysql/production/2026-09-29/Full_<id> && find . -type f -print0 \
  | LC_ALL=C sort -z | xargs -0 sha256sum | sed 's|  \./|  |' | sha256sum
```

Records taken before this release have no checksum. Record them explicitly:

```bash
bdbackup history checksum --config config.toml
bdbackup history checksum --config config.toml --job mysql-prod
```

It prints `<id>: recorded`, `<id>: skipped: unavailable` for an artifact that is
missing or was replaced, or `<id>: skipped: unexpected unit` for a physical backup
outside its dated layout, and then exits 1. A checksum already recorded is never
changed. The checksum lives in the `checksum` column of the `backup_runs` table.

History uses schema version 5. An existing database upgrades on the next write;
older bdbackup versions refuse a version-5 database.

History is written before work starts, and success only after the backup and its
requested verification finish. If history cannot be written, the command fails;
an artifact already created is preserved. A forcibly killed process may leave a
`running` row, which is never offered for restore. Existing backups are not
imported automatically.

Select a successful backup interactively:

```bash
bdbackup restore --config config.toml
bdbackup restore --config config.toml --job files-daily
```

Each backup job may set `restore_root = "/restore/mysql-prod"` to override
`[history] restore_root` for its guided-restore suggestions. It must be a nonempty
path string; `~` expands and relative paths resolve against the config file.
Job-level `restore_root` requires `[history]` and is not allowed on retention jobs.
Validation checks that each job's root is writable or can be created.
Physical jobs usually need a separate root with room for the full and its
incrementals, staged near the database datadir.

The current config's job is matched by the history record's job name. If that job
has no `restore_root`, or is no longer configured, the history root is used.
An explicit `--dst` always wins.

Restore lists successful, available artifacts, asks for the backup ID, suggests
`<restore_root>/<job>-<id>`, and asks for confirmation. You can edit the suggested
path. Repeated restores suggest a numbered alternative; history-based restore
requires a new directory even when `--dst` is supplied. For automation, specify
the exact backup ID and use `--yes`:

```bash
bdbackup restore --config config.toml --backup-id 12 --dst /restore/job-12 --yes
```

- File archives are extracted into the selected directory with the existing
  safe extraction filters.
- MySQL dumps are decompressed and checked into `<destination>/backup.sql`.
  Importing SQL into a running server is a separate administrator action.
- XtraBackup/MariaDB backups produce a prepared recovery directory, including
  prerequisite increments. The destination must be outside the backup root.
  Server ownership, copy-back, and startup remain administrator actions.
- Custom engines are recorded, but require their own restore support unless
  they inherit a supported backend.

Deleted artifacts remain in history as unavailable. File identity, size, and
modification time detect replaced archives, so older rows for a reused filename
are not offered as older recovery points. Availability uses these file checks.
Restore checks the recorded checksum of every file it uses for backups that GFS
manages (see [Ledger and GFS](#ledger-and-gfs)); for other backups it
validates the actual archive or physical dependency chain. History stores
references to artifacts and does not preserve an archive that a later backup
replaces.

## Ledger and GFS

GFS moves bdbackup's own backups through a chain of storage tiers, called stages,
using only the history ledger: it never guesses from what lies on disk. A typical
chain keeps everything for 5 days on a fast local disk, everything for 15 more
days on cheap NFS storage (point-in-time recovery from incrementals), then one
backup per week, per month and per year.

```toml
[gfs-main]
type = "gfs"
apply = false               # dry run unless true
# schedule = "0 4 * * *"

[[gfs-main.stage]]
paths = ["/BACKUP"]         # engines write here
period = "daily"
keep = "5d"

[[gfs-main.stage]]
paths = ["/NFS/daily"]
period = "daily"
keep = "20d"                # 5 days hot + 15 days on NFS

[[gfs-main.stage]]
paths = ["/NFS/weekly"]
period = "weekly"
keep = "8w"

[[gfs-main.stage]]
paths = ["/NFS/monthly"]
period = "monthly"
keep = "12m"

[[gfs-main.stage]]
paths = ["/NFS/yearly", "/DD/yearly"]   # every path receives a copy
period = "yearly"
keep = "7y"
```

Run it with `bdbackup run --config config.toml gfs-main`. `--cron` schedules GFS
jobs with retention at 04:00, after backups.

**Stages and ages.** `keep` is an age counted from the backup day: `Nd` days,
`Nw` weeks, `Nm` calendar months, `Ny` years. A backup belongs to the first stage
whose `keep` it is still within; past the last stage it is deleted. Each `keep`
must be longer than the previous one on every calendar (a month counts as 28–31
days), and periods never go backwards along the chain. The first stage has
exactly one path and is where backup jobs write.

**Periods.** A daily stage keeps every backup. A weekly stage keeps the newest
backup of each ISO week; a monthly stage, of those, the newest dated in each
month; a yearly stage, of those, the newest dated in each year. Selection is per
job, so a tar job and a MySQL job never compete for one slot. A week is decided
only once it has ended, and a month or year only once the ISO week holding its
last day has ended, so a choice never changes later; until then the backup is
reported `held: <bucket> not complete` and stays where it is.

With the configuration above on Tuesday 2026-09-29 and one backup per day:

| Backup | Result | Why |
|---|---|---|
| 2026-09-24 | `/NFS/daily` | within 20d |
| 2026-09-09 | deleted | its week's newest backup is 09-13 |
| 2026-09-06 (Sun) | `/NFS/weekly` | newest of ISO week 36 |
| 2026-08-30 (Sun) | `/NFS/weekly` | newest of week 35; also August's monthly backup |
| 2026-08-02 (Sun) | deleted | newest of its week, not of August |
| 2026-07-26 (Sun) | `/NFS/monthly` | July's newest weekly backup |
| 2025-12-28 (Sun) | `/NFS/monthly` | also 2025's yearly backup |
| 2024-12-29 (Sun) | `/NFS/yearly` | 2024's yearly backup |

**Units and paths.** GFS moves a unit as one piece: an archive or dump file, or
an xtrabackup dated directory with its full and every incremental (and its
`Full_Latest` and `.full_success` markers). A unit keeps its path relative to
the first stage, so `/BACKUP/mysql/prod/2026-09-29` becomes
`/NFS/daily/mysql/prod/2026-09-29`. The newest unit of each job always stays in
the first stage, so the next incremental finds its full.

**Safety.** Each later stage path must contain an empty file named
`.bdbackup-destination`; an unmounted mount point is an empty local directory
without it, and GFS writes nothing there (`--validate` reports it). GFS copies to
a temporary name, checks every checksum while reading the source, and removes
the original only after every copy is in place and recorded; a failed copy
removes its temporary copies and the next run retries. Before a move,
every copy of the unit is checked, not only the one copied, and
copies keep the permission bits of every file and directory. If the
ledger lists a unit in two stages (left by an earlier version or a hand edit),
the next run removes the earlier copy only after the later copies verify.
Copies hold no backup lock,
so a slow NFS copy never blocks a backup; if an xtrabackup root is locked when
GFS commits, the unit is reported `deferred: locked` and retried next run. GFS
only touches backups recorded in the ledger. It refuses, and reports:

- `refused: checksum mismatch`: the backup changed since it was recorded. To
  accept the change, set the `checksum` column of its `backup_runs` rows to the
  new value with `sqlite3`; the next run acts on it. A backup that changed
  before GFS first saw it is never managed.
- `refused: unexpected entry …`: a unit holds something that is not one of its
  recorded backups or engine markers, such as a leftover `.tmp` directory, or a
  symlink or special file inside a backup.
- `refused: stage not configured`: the unit lives under a stage no longer in the
  configuration. Removing a stage never deletes its backups; restore the stage
  or remove them yourself. Changing the first stage's path keeps managing units
  already moved to later stages; units left under the
  old first-stage path are no longer managed.
- `refused: destination not ready …`: a stage path lacks its marker.

Backup jobs writing into the first stage must not delete or overwrite on their
own: xtrabackup jobs set `retention_days = 0` and file jobs set
`timestamp = true`. Configuration validation enforces both, rejects stage paths
that overlap, and keeps the live history database out of stage paths.

`apply = false` (the default) prints what would move or be deleted and changes
nothing. Exit codes: 0 done, 1 a refusal or failure, 2 configuration error, 3
another run of the same job is in progress, or the only unfinished units were
deferred.

**Unmanaged backups.** A backup that was unavailable when GFS first saw it (for
example an older run of a file job without `timestamp`, whose archive a later
run replaced) stays unmanaged on every later run: GFS never moves or deletes it,
and the report counts it in `unmanaged`.

**Report.** Each run lists every unit it acted on, held, deferred or refused,
then a `Summary:` line with counts and bytes per action. One line per stage
follows, such as `Stage /NFS/daily: 15 units (3200000000 bytes)`, counting where
the units are when the run ends (where they would be, in a dry run), and then
the ledger copies written.

**Step log.** Every step is also recorded in the `gfs_steps` table of the history
database, one row per path, with its source, destination, outcome
and reason. A failed step keeps the error message.

**Run log.** Every applying run records one row in the `gfs_runs` table:
the GFS job, start and finish time, status, exit code and the Summary counts.
The status is `running` while the run works, `completed` once it reaches its report
(whatever the exit code), and `failed` if the run itself crashed. A row left
`running` shows a run that was cut off; check that run's steps for leftovers
(Known limits below). Each `gfs_steps` row names its run in `run_id`. A dry run
records nothing.

**Running backups.** While a job has a backup in state `running` in the
ledger, GFS leaves every backup of that job where it is, whatever its age, and
reports each one as
`deferred: backup running (record <id>, started <time>)`.
Other jobs proceed as usual, and a run whose only unfinished work is these
deferrals exits 3. A backup killed before it finished (`kill -9`, a power loss)
leaves its row `running`, and that job stays deferred until you fix the row. After
checking that no backup of that job is running, mark it failed:
`sqlite3 <history database> "UPDATE backup_runs SET status = 'failed' WHERE id = <id>"`.

**Restoring moved backups.** `bdbackup restore` and `bdbackup history` find a
backup where GFS put it. `history` shows a copy that was removed or replaced as
`unavailable`. Before restoring, every backup file used is checked against its
ledger checksum; with several paths in a stage, the first path that verifies is
used, and the `Backup:` line names it. An xtrabackup chain is prepared from the
stage directory holding it, which keeps the layout of the first stage.

**Ledger copies.** After every applying run, each path of every stage after the
first holds `<gfs job>.ledger.sqlite3`, a consistent copy of the whole history
database, written under a temporary name and renamed into place. A path without
its marker gets none, and the run exits 1. To restore with only a stage left,
copy that file somewhere outside the stages, point `[history] database` at the
copy, and run `bdbackup restore`. Paths in the ledger are absolute, so the
stages must be mounted at the same paths as when the copy was written.

**Known limits.** GFS does not yet keep a journal of a move or delete while it
runs, so a few failures leave work for the operator. The report and the step log
name the paths involved, and each leftover is
removed by hand:

- A move was recorded, but removing the old copy failed. The old copy stays on
  disk, no longer in the ledger, and GFS never touches it again. Delete it.
- A move failed, or the process stopped, after some new copies were renamed into
  place but before the move was recorded. The next run reports
  `refused: final name exists: <path>`. Delete that copy; the next run moves the
  unit again.
- A delete failed part way. The next run reports `refused: missing location …`
  for a copy already removed. Delete its row with
  `sqlite3 <history database> "DELETE FROM gfs_locations WHERE path = '<path>'"`.
  If that was the unit's last row, also mark the unit deleted, or later runs
  report `refused: stage not configured`:
  `sqlite3 <history database> "UPDATE gfs_units SET deleted_at = datetime('now') WHERE id = <unit_id>"`
  (the `unit_id` of the row you deleted).
- A restore stops when the xtrabackup copy at the first path of a stage has
  damaged checkpoints or chain metadata; it does not move on to the next path.
  Prepare the copy at another path of that stage with
  `bdbackup xtrabackup prepare -r <directory holding the copy> -d <new dir> <backup>`.

## Development

Run tests and linting:

```bash
pytest
ruff check bdbackup tests scripts
```

Optional database integration tests use disposable, network-isolated Docker
containers with synthetic data. They remove only the containers and volumes
created for that run. Pull the matching images before running:

```bash
python scripts/integration_mysql.py           # mysql:8.4
python scripts/integration_physical.py percona # percona-server/xtrabackup:8.4
python scripts/integration_physical.py percona-encrypted # AES256 + zstd recovery
python scripts/integration_physical.py mariadb # mariadb:11.4
```

Physical recovery returns a prepared data directory; copying it into a server's
data directory, assigning ownership to the server user, and starting that server
are separate administrator actions. Use the same database version and
filesystem case-sensitivity as the source.

`pyproject.toml` (`project.version`) is the sole version source. The package's
`__version__` and CLI read installed distribution metadata generated from it.
After changing the version, run `uv lock` to refresh the generated lockfile and
reinstall with `pip install -e '.[dev]'` to refresh local metadata. Source
development requires this editable installation.

Build and validate a release:

```bash
python -m build
twine check dist/*
python scripts/smoke_install.py dist
```

Commit the validated changes and create an annotated tag named
`v<project.version>`. The publishing workflow accepts only that matching tag. It
runs the reusable CI workflow on that commit (tests on Python 3.12–3.14, lint,
real MySQL/Percona/MariaDB recovery tests, source/wheel build, Twine, and an
installed-wheel recovery smoke test), then
uploads those exact artifacts. Configure the PyPI trusted publisher for
`ayder/bdbackup`, `publish.yml`, environment `pypi` before releasing. Account
configuration and actual database recovery evidence are release prerequisites.

## License

MIT License. See [LICENSE](https://github.com/ayder/bdbackup/blob/main/LICENSE).
