Metadata-Version: 2.5
Name: shockbench-flow
Version: 0.1.2
Summary: Benchmark for online decision algorithms on a dynamic supply graph under calibrated, partially observable disruptions
Author-email: nuinashco <nuinashco@users.noreply.github.com>
License-Expression: MIT
License-File: LICENSE
License-File: THIRD_PARTY_NOTICES
Keywords: benchmark,gymnasium,network-flow,online-algorithms,operations-research,reinforcement-learning,supply-chain
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: <3.14,>=3.12
Requires-Dist: fastjsonschema~=2.22.2
Requires-Dist: highspy==1.12.0
Requires-Dist: joblib~=1.5.2
Requires-Dist: loguru~=0.7.3
Requires-Dist: networkx~=3.7
Requires-Dist: numba==0.67.0
Requires-Dist: numpy==2.4.5
Requires-Dist: scipy==1.18.1
Provides-Extra: gym
Requires-Dist: gymnasium~=1.3.0; extra == 'gym'
Requires-Dist: matplotlib~=3.11.2; extra == 'gym'
Provides-Extra: rl
Requires-Dist: gymnasium~=1.3.0; extra == 'rl'
Requires-Dist: matplotlib~=3.11.2; extra == 'rl'
Requires-Dist: stable-baselines3~=2.9; extra == 'rl'
Requires-Dist: torch<3,>=2.6; extra == 'rl'
Description-Content-Type: text/markdown

# shockbench-flow

A benchmark for online decision algorithms on a supply network under disruption,
packaged as a Gymnasium task.

Every week of an episode an agent decides how much of each good to send along
each route of a network (fuel to power grids, wafers to chip factories, chips to
markets). Before the episode starts, disruptions are drawn at random and nothing
the agent does changes them: a sea strait closes, a route is sanctioned, a
tariff jumps, a factory goes down. The agent sees the network as it is this
week, its stock and shipments, a demand forecast and noisy early warnings. An
episode costs money (freight, tariffs, storage, unmet demand); lower is better.

The **score** compares the agent's cost with two references on the same
scenarios: **0** is a simple rule that keeps shipping the normal plan and
ignores disruptions (the naive rule), **1** is a clairvoyant plan that knew
every disruption in advance. Below 0 is worse than the naive rule.

## Install

Python 3.12 or 3.13:

```bash
pip install shockbench-flow            # the simulator, the scoring and the agent kit
pip install "shockbench-flow[gym]"     # and the Gymnasium environments
pip install "shockbench-flow[rl]"      # and Stable-Baselines3 and PyTorch, for PPO
```

On Linux, add `--extra-index-url https://download.pytorch.org/whl/cpu` to the
`rl` install for PyTorch's CPU build (PyPI's Linux wheel bundles CUDA). NumPy,
SciPy, HiGHS and Numba are pinned to exact versions, since a replay is
bit-identical only on the same numerical libraries.

Hackathon participants install it through the starter repository they were
given, which pins the version the server scores with (`uv sync` there).

The wheel installs three packages: `shockbench_flow` (the simulator and the
scoring), `shockbench_flow_gym` (the Gymnasium environments, extra `gym`) and
`shockbench_flow_agent` (the agent kit: the submission contract, the local
evaluation and the scoring API).

## Environments

```python
import gymnasium as gym
import shockbench_flow_gym  # registers the ShockBench/* environments

env = gym.make("ShockBench/Tiny-v0")  # also ShockBench/Small-v0 and ShockBench/Full-v0
obs, info = env.reset(seed=0)
```

| Environment           | Network                       |
| --------------------- | ----------------------------- |
| `ShockBench/Tiny-v0`  | a small network to start with |
| `ShockBench/Small-v0` | the Development phase network |
| `ShockBench/Full-v0`  | the Final phase network       |

The observation and the action are dicts of numpy arrays with fixed shapes; the
reward is minus the week's cost in dollars.

## The Agent contract

A submission is a zip with `agent.py` at its root and any files it needs beside
it:

```python
class Agent:
    def __init__(self, config):
        ...  # once per episode: config holds the public tables, the horizon T, policy_seed and the array shapes

    def act(self, observation):
        ...  # once per week: observation is the Gymnasium Dict observation
        return {"flows": flows, "override_qty": override_qty, "release_mode": release_mode}
```

A week whose `act` raises or returns a malformed action is played by the naive
rule, and counted.

- `python -m shockbench_flow_agent.submission submission.zip` checks a zip as
  the server does, without running it.
- `shockbench_flow_agent.evaluate("submission.zip")` (or a folder, an
  `agent.py`, an `Agent` class) scores it on the public dev episodes, and
  `shockbench_flow_agent.report(summary)` prints a plain report: the score, its
  interval over the choice of episodes, the costs in dollars and the weeks the
  naive rule played.

For search loops, `EpisodeSet` computes the reference costs once and caches them
on disk:

```python
from shockbench_flow_agent import EpisodeSet

episodes = EpisodeSet.build("tiny", "dev")  # the leaderboard's local dev split
result = episodes.score(Agent)              # an Agent class, a factory, a folder, a zip or an agent.py
print(result)                               # score, interval, costs in dollars, fallback weeks
print(episodes.compare(Agent, OtherAgent))  # paired comparison on the same episodes
EpisodeSet.build("tiny", "dev", quick=True) # seconds instead of minutes, not the leaderboard's numbers
```

Under Gymnasium, `shockbench_flow_gym.agent_config_from_reset(env, obs, info)`
is the `config` the scorer builds for the episode a reset started, and
`shockbench_flow_gym.play_episode(env, Agent, episode)` plays one episode and
returns its cost in dollars.

## Limits

`shockbench_flow_agent.LIMITS` holds what the server applies to every
submission:

| Limit               | Value                                                                                        |
| ------------------- | -------------------------------------------------------------------------------------------- |
| wall clock per week | 10 s for `act`'s reply (`deadline_s`)                                                        |
| CPU per week        | 2 s on `small` (Development), 4 s on `full` (Final) (`cpu_budget_s`)                         |
| start-up            | 60 s to start and import `agent.py`, charged to no week (`startup_s`)                        |
| one episode         | stopped 192.9 s (`small`) or 500 s (`full`) after its container starts (`episode_timeout_s`) |
| a reply             | at most 1 MiB (`max_reply_bytes`)                                                            |
| the container       | 1 CPU, 4 GB, 128 processes, no network, a read-only file system (`container`)                |
| the submission      | 500 MiB unpacked, 1,000 files (`submission`)                                                 |

A week over its wall clock or its CPU budget is played by the naive rule;
`Agent(config)` counts toward week 1. `play_isolated` times each week of an
episode in a process that has only the scoring image's packages, and
`play_container` plays it in a local copy of the scoring container
(`build_image`) under the server's CPU meter.

## Licence

MIT. `THIRD_PARTY_NOTICES` in the distribution carries the notices of the
NumPy-derived code and the attributions of the data sources behind the packaged
networks.
