Metadata-Version: 2.5
Name: npflow
Version: 0.2.0
Summary: Neural Petri Flow. Neural networks that remain Petri nets for every value of their weights
Project-URL: Homepage, https://github.com/daenuprobst/npf
Project-URL: Repository, https://github.com/daenuprobst/npf
Author: Daniel Probst
License-Expression: MIT
License-File: LICENSE
Keywords: atom-mapping,chemistry,graph-neural-network,petri-net,rdkit,reaction-prediction
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Chemistry
Requires-Python: <3.14,>=3.12
Requires-Dist: lightning>=2.6.6
Requires-Dist: networkx>=3.2
Requires-Dist: numpy>=2.0
Requires-Dist: rdkit>=2026.3.6
Requires-Dist: scipy>=1.13
Requires-Dist: torch>=2.6
Provides-Extra: cpsat
Requires-Dist: ortools>=9.15.6755; extra == 'cpsat'
Description-Content-Type: text/markdown

# Neural Petri Flow

A neural network that is a Petri net rather than one that runs on a Petri net. The learned parts are the rate law and,
for classification, a readout of the firing. Conservation, enabling and the state equation are parameter-free layers,
so for every value of the weights the outputs are markings of the given net and firing counts that satisfy its state
equation.

A chemical reaction is written as a firing sequence of a valence net. Bond places hold bond orders, slack places hold
free valence, transitions make and break bonds. Forward prediction, atom mapping and reaction classification become
three questions about one firing vector.

## Install

From PyPI, for Python 3.12 and 3.13:

    pip install npflow                # or: uv add npflow
    pip install "npflow[cpsat]"       # also the integer program of the first mapper, which needs OR-tools

The atom mapper needs no training and runs from the terminal on one reaction or on a file with one per line, an
optional identifier after a tab. A reaction that cannot be mapped gives an empty line, so the lines stay in order.

    npflow map "CC(=O)Cl.CN>>CC(=O)NC"
    npflow map reactions.smi -o mapped.smi --processes 8

In Python, `from npflow.chem import map_reaction`, and the models are in `npflow.chem.api` (see Your own reactions). To
work on the code or reproduce the paper, clone the repository, which holds the benchmarks, their results and the tests:

    uv sync

## Layout

Models and data live in the package, one model per file. Everything that produces a number of the paper lives under
`benchmarks/` and is run as a module.

### Package

| File | What it does |
|---|---|
| `src/npflow/nets.py` | attributed nets, incidence matrix, invariants, random and chain nets |
| `src/npflow/simulate.py` | stochastic token game and continuous flow |
| `src/npflow/datasets.py` | pairs of markings with firing counts, next-state samples |
| `src/npflow/batching.py` | many nets as one disjoint net |
| `src/npflow/layers.py` | enabling, conflict resolution, token game flow, state-equation projection |
| `src/npflow/models/npf.py` | rate law, token game, bridge, projection, occupancy refinement |
| `src/npflow/models/pgnn.py` | PGNN as published, and its variants |
| `src/npflow/models/linear_pgnn.py` | the linear special case of the PGNN paper |
| `src/npflow/models/gnn.py` | message passing on the place graph |
| `src/npflow/models/state_equation_only.py` | projection without learning |
| `src/npflow/models/equilibrium.py` | equilibrium layer, learned place energies, dual Newton |
| `src/npflow/models/learned_incidence.py` | learned arc weights |
| `src/npflow/models/coloured.py` | coloured tokens |
| `src/npflow/chem/featurisation.py` | reactions as markings of the valence net |
| `src/npflow/chem/encoder.py` | message passing on the valence net |
| `src/npflow/chem/decode.py` | products, by-products and atom maps read from a marking |
| `src/npflow/chem/token_game.py` | forward prediction as a firing sequence with enabling |
| `src/npflow/chem/one_shot.py` | one-shot counterpart of the token game |
| `src/npflow/chem/classifier.py` | state-equation readout, and the explicit firing vector |
| `src/npflow/chem/targets.py` | training targets from the net, and the scoring of a marking against the recorded product |
| `src/npflow/chem/mapper.py` | the atom mapper, the minimum firing vector without a solver, all its ties, and the third level that chooses among them |
| `src/npflow/chem/search.py` | the branch and bound behind the mapper, bounded by linear assignments on the places of the net |
| `src/npflow/chem/third_level.py` | the three counts that rank maps of equal cost, oxidation state of carbon, aromatic bonds, electron sinks |
| `src/npflow/chem/cost.py` | the two levels of the cost of a mapping |
| `src/npflow/chem/cgr.py` | the condensed graph of reaction, which merges maps into classes and scores a map against a reference |
| `src/npflow/chem/mapping.py` | the maps of the mapper as SMILES with map numbers |
| `src/npflow/chem/api.py` | train, validate, test and use the models on your own reactions, with PyTorch Lightning |
| `src/npflow/chem/exact.py` | the same minimum as an integer program with CP-SAT, the first version of the mapper, kept as the reference the search is checked against |
| `src/npflow/chem/open_net.py` | source transitions for reactions whose reactants the record omits |
| `src/npflow/chem/minimise.py` | a heuristic seating, the start of the integer program |
| `src/npflow/chem/orders.py` | enabled linearisations of a firing vector |
| `src/npflow/chem/verifier.py` | re-ranker for token game candidates, no gain, kept for the record |
| `src/npflow/chem/arrows.py` | a mechanistic step as a net of electron pairs, arrows as transitions, the octet rule as enabling |
| `src/npflow/chem/arrow_game.py` | elementary steps as firing sequences of that net |

### Benchmarks

| File | What it does |
|---|---|
| `benchmarks/synthetic/experiment.py` | one run on one regime |
| `benchmarks/synthetic/sweep.py` | all models on one task |
| `benchmarks/synthetic/paper_protocol.py` | the protocol of the PGNN paper |
| `benchmarks/synthetic/locality.py` | chain nets of growing length |
| `benchmarks/synthetic/equilibrium.py` | equilibria of unseen reversible nets |
| `benchmarks/synthetic/learned_incidence.py`, `coloured.py` | extensions |
| `benchmarks/chemistry/build_data.py` | Schneider 50k and USPTO-MIT to `data/` |
| `benchmarks/chemistry/experiment.py` | forward and classify, all datasets |
| `benchmarks/chemistry/care.py` | EC number classification on CARE task 2 and on ECREACT in the Enzyformer split, `--maps` adds the firing vector of a map file |
| `benchmarks/chemistry/enzymes.py` | the enzymatic data, EnzymeMap at EC level 3 and ECREACT in the Enzyformer split, and the atom maps of RXNMapper and of the record |
| `benchmarks/chemistry/exact_map.py` | the mapper on a data set, and the maps the classifier reads |
| `benchmarks/chemistry/exact_map_report.py` | comparison with RXNMapper, intervals, sign test |
| `benchmarks/chemistry/synrxn_map.py` | the learning-free mapper on the five SynRXN sets, scored with SynKit like the published mappers |
| `benchmarks/chemistry/synrxn_enzymemap.py` | the mapper and RXNMapper on EnzymeMap, scored against its curated maps with the validator of SynRXN |
| `benchmarks/chemistry/mechanism.py` | elementary steps of the FlowER mechanism benchmark, `--net arrow` or `--net electron` |
| `benchmarks/chemistry/mechanism_validity.py` | untrained arrow and electron nets, with and without the octet rule, end only in valid molecules |
| `benchmarks/chemistry/mechanism_pathways.py` | top-k pathway accuracy on FlowER from the saved per-step ranks, FlowER's metric |
| `benchmarks/published/` | the published numbers the paper compares with, each row with its source |
| `benchmarks/chemistry/exact_ties.py` | the mappings the net cannot tell apart |
| `benchmarks/chemistry/balance.py` | balanced equations from the open net |
| `benchmarks/chemistry/golden.py` | the Golden atom mapping set |
| `benchmarks/chemistry/errors.py` | where the token game fails |
| `benchmarks/chemistry/decoding_rules.py`, `calibrate.py`, `ensemble.py`, `verify.py` | decoding and ranking |
| `benchmarks/chemistry/net_targets.py` | forward training targets from the net, every minimum firing vector that reaches the product |
| `benchmarks/chemistry/insights.py`, `invariance.py`, `figures.py` | analyses and figures |
| `benchmarks/chemistry/baselines/molecular_transformer.py` | the baseline, trained here |
| `benchmarks/checks/paper_checks.py` | numerical check of every proposition, exits non-zero on failure |
| `benchmarks/checks/data_facts.py` | the facts about the data that the paper quotes |
| `benchmarks/report.py` | `results/*.json` to `results/REPORT.md` |
| `benchmarks/paper_tables.py` | `results/REPORT.md` to `paper_v2/tables/appendix_tables.tex` (chemistry) and `paper_v2/tables/synthetic_tables.tex` |
| `tests/` | unit tests and end-to-end runs of the experiment scripts |

## Reproduce

Where each benchmark of the paper comes from. The commands are in the paragraphs below, and `benchmarks.report` gathers
the result files into the tables of the paper.

| Benchmark | Script | Results |
|---|---|---|
| Atom mapping, Golden set | `exact_map`, `exact_map_report` | `results/chem/exact_map/` |
| Atom mapping, SynRXN sets | `synrxn_map` | `results/chem/synrxn_map/` |
| Atom mapping, EnzymeMap | `enzymes`, `exact_map`, `synrxn_enzymemap` | `results/enzymemap_ec/synrxn_map.json` |
| Reaction classes, Schneider 50k | `experiment --task classify` | `results/chem/classify/` |
| EC numbers, CARE task 2 | `care` | `results/care/easy/` |
| EC numbers, ECREACT (Enzyformer split) | `enzymes`, `care --dataset ecreact_enzyformer` | `results/ecreact/enzyformer/` |
| EC numbers, EnzymeMap by map source | `enzymes`, `experiment --dataset enzymemap_ec` | `results/enzymemap_ec/classify/` |
| Forward prediction, USPTO-480K, and its validity | `experiment --task forward --dataset uspto_mit` | `results/uspto_mit/forward/` |
| Elementary steps, FlowER | `mechanism` | `results/mechanism/npf*-[0-2].json` |
| Pathways, FlowER | `mechanism_pathways` | `results/mechanism/pathways.json` |
| Validity for all weights, mechanism nets | `mechanism_validity` | `results/mechanism/validity.json` |
| Published numbers compared with | | `benchmarks/published/` |

Tests, the propositions checked numerically, and the facts about the data that the text quotes.

    uv run pytest -q
    uv run python -m benchmarks.checks.paper_checks
    uv run python -m benchmarks.checks.data_facts

Synthetic nets (Appendix E). The `npf-kl` rows come from `benchmarks.synthetic.experiment` with `--model npf-kl` on every
regime of the sweep.

    uv run python -m benchmarks.synthetic.sweep --task transitions --workers 10
    uv run python -m benchmarks.synthetic.sweep --task next --iters 3000 --workers 10
    uv run python -m benchmarks.synthetic.paper_protocol
    uv run python -m benchmarks.synthetic.locality
    uv run python -m benchmarks.synthetic.equilibrium
    uv run python -m benchmarks.synthetic.learned_incidence
    uv run python -m benchmarks.synthetic.coloured

Data. The Golden set is the RDF of Lin et al. (2022).

    uv run python -m benchmarks.chemistry.build_data                    # data/schneider50k.pkl
    uv run python -m benchmarks.chemistry.build_data uspto-mit          # data/uspto_mit.pkl
    uv run python -m benchmarks.chemistry.golden prepare <golden.rdf>   # data/golden.pkl, data/golden_unmapped.txt

Atom mapping. The mapper of the paper on the Golden set and on EnzymeMap, with its third level, then the chosen cost and
its alternatives, the 200 dev reactions on which the cost was chosen, the ties, the five SynRXN sets, the open net, and
RXNMapper on the same reactions, which runs in an environment of its own. With the chosen cost `exact_map` runs the
mapper of the paper, with its third level, and `--no-third-level` keeps the first optimum. Add `--solver cp-sat` to
`exact_map` or `synrxn_map` for the integer program of the first version, the only path that loads OR-tools.

    uv run --no-project --python 3.11 --with rxnmapper --with rdkit --with "setuptools<81" --with "numpy<2" python benchmarks/chemistry/baselines/rxnmapper_golden.py
    uv run python -m benchmarks.chemistry.exact_map --data golden --third-level
    uv run python -m benchmarks.chemistry.exact_map --data enzymemap_3k --third-level
    uv run python -m benchmarks.chemistry.exact_map_report results/chem/exact_map/enzymemap_3k-1-1-1-nolabile-ch-third.pkl results/chem/rxnmapper_enzymemap_3k.json
    uv run python -m benchmarks.chemistry.exact_map --data golden-dev --labile-h --no-ch-places
    uv run python -m benchmarks.chemistry.exact_map --data golden-dev --no-ch-places
    uv run python -m benchmarks.chemistry.exact_map --data golden-dev --no-third-level
    uv run python -m benchmarks.chemistry.exact_map --data golden --secondary 0,0,0 --labile-h --no-ch-places
    uv run python -m benchmarks.chemistry.exact_map --data golden --labile-h --no-ch-places
    uv run python -m benchmarks.chemistry.exact_map --data golden --no-ch-places
    uv run python -m benchmarks.chemistry.exact_map --data golden --no-third-level
    uv run python -m benchmarks.chemistry.exact_ties
    uv run python -m benchmarks.chemistry.synrxn_map map --seconds 60
    uv run --with "synkit>=1.5,<1.6" python -m benchmarks.chemistry.synrxn_map score --seconds 60
    uv run python -m benchmarks.chemistry.balance --data golden
    uv run python -m benchmarks.chemistry.balance --data schneider50k
    uv run python -m benchmarks.chemistry.balance "CC(=O)Cl.NCCN>>CC(=O)NCCNC(C)=O"

Targets of the net. The maps that classification reads, and the firing vectors that forward prediction trains on.
`experiment --task forward` builds a missing target file itself, so these commands only build ahead of time. The
forward runs of the paper trained on the files of the integer program, kept as `data/net_targets_<name>-cpsat.pkl`,
pass them with `--net-targets` to reproduce those numbers.

    uv run python -m benchmarks.chemistry.exact_map --data schneider50k --third-level --processes 12 --write data/exact_maps_schneider50k.pkl
    uv run python -m benchmarks.chemistry.net_targets --dataset schneider50k
    uv run python -m benchmarks.chemistry.net_targets --dataset uspto_mit --subset 40900
    uv run python -m benchmarks.chemistry.net_targets --dataset uspto_mit

Classification on Schneider 50k. Seeds 0 to 4 for the three models of the main table, 0 to 2 elsewhere.

    C="uv run python -m benchmarks.chemistry.experiment --task classify"
    $C --model npf --seed 0                                          # also npf-sigma, pgnn-sigma, pgnn, npf-nogate, drfp
    $C --model npf --size-split --seed 0                             # also npf-nogate, pgnn
    $C --model npf --labels 1000 --select last --seed 0              # also 250 labels; pgnn, npf-sigma, pgnn-sigma, drfp
    $C --model npf --firing --labels 1000 --select last --seed 0     # also 250 labels; pgnn
    uv run python -m benchmarks.chemistry.ensemble npf-sigma         # also npf
    uv run python -m benchmarks.chemistry.invariance
    uv run python -m benchmarks.chemistry.insights

EC numbers of enzymatic reactions, seeds 0 to 2, with atom maps from three sources: the mapper of the paper, RXNMapper
and, on EnzymeMap, its curated maps. The same classifier reads the firing vector of each, against the same classifier
without maps. EnzymeMap at EC level 3 holds out 10 % of the reactions of every class and uses the state-equation readout
of Schneider 50k. CARE task 2 (easy split) and ECREACT in the split of Enzyformer (zenodo 18083829, downloaded and
checked by `prepare`) use the published CARE model, the readout with the participation gate. The model without maps on
EnzymeMap reads no map, whichever file `--maps` names. A reaction the mapper cannot prove optimal within 60 s keeps the best map
found so far, which depends on the machine: rebuilt on CARE, 5 of 55,081 maps differed, all of them unproved. The data,
the unmapped reactions for RXNMapper and the maps:

    uv run python -m benchmarks.chemistry.care prepare <path to CARE_datasets>
    uv run python -m benchmarks.chemistry.enzymes prepare enzymemap
    uv run python -m benchmarks.chemistry.enzymes prepare enzyformer
    for data in enzymemap_ec care_easy; do
      uv run python -m benchmarks.chemistry.exact_map --data $data --third-level --processes 16 --write data/enzyme_maps/mapper_maps_$data.pkl
      uv run python -m benchmarks.chemistry.enzymes unmapped $data
      uv run --no-project --python 3.11 --with rxnmapper --with rdkit --with "setuptools<81" --with "numpy<2" python benchmarks/chemistry/baselines/rxnmapper_golden.py data/enzyme_maps/${data}_unmapped.txt data/enzyme_maps/rxnmapper_$data.json
      uv run python -m benchmarks.chemistry.enzymes convert $data                # rxnmapper_maps_$data.pkl, recorded_maps_$data.pkl
    done

Classification, `results/enzymemap_ec/classify`, `results/care/easy` and `results/ecreact/enzyformer`:

    E="uv run python -m benchmarks.chemistry.experiment --task classify --dataset enzymemap_ec --batch 16"
    $E --model npf-nogate --maps data/enzyme_maps/mapper_maps_enzymemap_ec.pkl --tag=-none --seed 0
    $E --model npf-sigma --maps data/enzyme_maps/mapper_maps_enzymemap_ec.pkl --tag=-mapper --seed 0   # also rxnmapper, recorded
    C="uv run python -m benchmarks.chemistry.care train"
    $C --seed 0
    $C --maps data/enzyme_maps/mapper_maps_care_easy.pkl --tag=-mapper --seed 0                         # also rxnmapper
    $C --dataset ecreact_enzyformer --seed 0

Atom mapping on EnzymeMap, the maps above against the curated maps of the 41,510 reactions that have one for every
product atom, scored with the validator of SynRXN and an exact sign test between the mappers.

    uv run --with "synkit>=1.5,<1.6" python -m benchmarks.chemistry.synrxn_enzymemap   # results/enzymemap_ec/synrxn_map.json

Forward prediction on Schneider 50k, seeds 0 to 2.

    F="uv run python -m benchmarks.chemistry.experiment --task forward"
    $F --model npf --seed 0                       # also npf-noenabling, npf-oneshot --matched, pgnn --matched
    $F --model npf --limit 2000 --seed 0          # also 8000; npf-noenabling, pgnn --matched
    uv run python -m benchmarks.chemistry.beam_width
    uv run python -m benchmarks.chemistry.figures tokengame 14440
    uv run python -m benchmarks.chemistry.figures attribution
    uv run python -m benchmarks.chemistry.figures loadbearing

Forward prediction on USPTO-MIT, seeds 0 to 2, and the Molecular Transformer baseline on the same subsets.

The full split, USPTO-480K, is run once per arm: the mapper's firing vectors as targets with and without the enabling
rule, and the recorded maps. The targets are `data/net_targets_uspto_mit.pkl`, built by the search, SHA-256
`b22581ce23dc4f52f6a2980b40e2ce3364fe0a6029e44f0be76507e14d9c060b`, which both of their result files record. The paper
reports `product_top{1,3,5}_beam_official`, over all 40,000 test lines, so the 6 that RDKit cannot parse count as wrong,
and `valence_valid`, the share of predictions whose touched fragments RDKit sanitises. The recorded-maps run was trained
before `valence_valid` was measured with RDKit and scored again with `--evaluate-only`, which leaves its accuracy and
beams unchanged.

    M="uv run python -m benchmarks.chemistry.experiment --task forward --dataset uspto_mit --amp"
    $M --model npf --width 256 --rounds 8 --attention 8 --lr 4e-4 --tag=-deep --seed 0
    $M --model npf-noenabling --width 256 --rounds 8 --attention 8 --lr 4e-4 --tag=-deep --seed 0
    $M --model npf --width 256 --rounds 8 --attention 8 --lr 4e-4 --tag=-deep --recorded-maps --seed 0
    M="$M --model npf"
    $M --subset 40900 --seed 0                                            # also --single-target, --subset 4090, --recorded-maps
    uv run python -m benchmarks.chemistry.decoding_rules results/uspto_mit/forward/npf-deep-nettargets-0.pt --dataset uspto_mit --width 256 --rounds 8 --attention 8
    uv run python -m benchmarks.chemistry.baselines.molecular_transformer prepare 40900    # then train 40900 --steps 30000 and score 40900; also 4090

Elementary steps on the FlowER mechanism benchmark, seeds 0 to 2.

Both models are token games. Given the reactants of one elementary step they fire arrows until STOP, and the marking
reached is the predicted products. The arrow net moves an electron pair per firing, so its transitions are the curly
arrows of arrow pushing and the octet rule is what enables them. The electron net moves one electron, so a fishhook is
a transition of weight one, radical steps become expressible, and the arrow net is its sub-net of weight two.
Molecules are Kekulé structures: tokens are electrons, and an aromatic bond of order 3/2 would hold three, half a pair.
Products are compared after aromaticity is perceived again, so the Kekulé structure chosen does not change the score.

The download stage fetches the published split from figshare and checks it, so a fresh machine needs nothing else.
Each net keeps its own prepared data, weights and results. The results in `results/mechanism` are seeds 0 to 2 of both
nets, seed 0 evaluated with a beam of 10 and seeds 1 and 2 with a beam of 5.

    uv run python -m benchmarks.chemistry.mechanism download
    uv run python -m benchmarks.chemistry.mechanism prepare --net electron --processes 20
    for seed in 0 1 2; do
      uv run python -m benchmarks.chemistry.mechanism train --net electron --seed $seed --epochs 12 --budget 2000000
      uv run python -m benchmarks.chemistry.mechanism evaluate --net electron --seed $seed --beam 10
    done
    uv run python -m benchmarks.chemistry.mechanism prepare --processes 20          # the arrow net, the same stages

Validity for all weights, untrained games of both nets with and without the octet rule as enabling, on 1,000 test
steps and 3 seeds.

    uv run python -m benchmarks.chemistry.mechanism_validity                      # results/mechanism/validity.json

Pathway accuracy, the metric of FlowER's `sequence_evaluation.py`: a test reaction counts at top k if some route from
its reactants to a terminal product has every step within the top k. It reads the per-step ranks that `evaluate` saves
next to each result, `results/mechanism/<name>-<seed>-ranks.npz`. FlowER's numbers, step and pathway accuracy and
validity, are the Source Data of its Figure 2, in `benchmarks/published/flower_fig2.csv`.

    uv run python -m benchmarks.chemistry.mechanism_pathways                      # results/mechanism/pathways.json

The report and the tables of the paper. Its last section, the benchmarks of the paper, puts NPF next to the published
numbers in `benchmarks/published`, each file with the paper and table it comes from.

    uv run python -m benchmarks.report > results/REPORT.md && uv run python -m benchmarks.paper_tables

## Your own reactions

The models train on reaction SMILES without atom maps. A file holds one reaction per line, `precursors>>product`, with
an optional tab separated class label. The mapper computes the targets once per file, milliseconds per reaction for
most reactions, and caches them next to it.

    from npflow.chem import api

    data = api.ReactionData("train.txt", "val.txt", "test.txt")
    model = api.ForwardModel()                                 # width=256, rounds=8, attention=8 is the USPTO-MIT model
    trainer = api.trainer(model, epochs=60)
    trainer.fit(model, data)
    trainer.test(model, data, ckpt_path="best")
    api.predict(model, ["CC(=O)Cl.NCC"])               # ranked products with probabilities
    api.map_reaction("CC(=O)Cl.NCC>>CC(=O)NCC")        # the reaction with atom map numbers

    data = api.ReactionData("train.txt", "val.txt", "test.txt", task="classify")
    model = api.ClassifierModel(data.n_classes, sigma=True)   # the classifier of the paper, reads the mapper's firing vector
    api.classify(model, ["CC(=O)Cl.NCC>>CC(=O)NCC"], data.classes)

`api.trainer` is a Lightning trainer with the settings of the paper, AdamW, a one-cycle schedule, gradient clipping and
the best epoch kept. Any Lightning trainer works, and `ForwardModel.load_from_checkpoint` reloads a run. Without
`sigma` the classifier is the state-equation readout, which needs no mapper at test time. `map_reaction` and the
training targets use the mapper of the paper, `npflow.chem.mapper`, the minimum firing vector of the chosen cost, found and
proved by a branch and bound without a solver, with a third level that chooses among its ties. The integer program of
`npflow.chem.exact` gives the same minimum and stays as a reference, and the benchmark scripts run it with `--solver cp-sat`.

## Conventions

Results are JSON under `results/`, one file per run, and every table of the paper is generated from them. No number
in the paper is typed by hand. Side outputs go in their own folder, because `benchmarks/report.py` globs the result
folders.

## Changelog

[anonymised]
