Metadata-Version: 2.5
Name: marl-battlegrounds
Version: 1.0.0
Summary: The JAX-native Benchmark for Heterogeneous and Competitive Multi-agent Reinforcement Learning.
Project-URL: Homepage, https://marl-battlegrounds.com
Project-URL: Documentation, https://marl-battlegrounds.com/#guides
Project-URL: Repository, https://github.com/oceansystemslab/MARL-BattleGrounds
Project-URL: Issues, https://github.com/oceansystemslab/MARL-BattleGrounds/issues
Project-URL: Changelog, https://github.com/oceansystemslab/MARL-BattleGrounds/releases
Author-email: Ulixes Tariq Hawili <t.hawili@sms.ed.ac.uk>
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: JAX,benchmark,competitive games,heterogeneous agents,multi-agent reinforcement learning,reinforcement learning,team deathmatch
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.12
Requires-Dist: chex
Requires-Dist: flatbuffers>=24
Requires-Dist: huggingface-hub<3,>=2.2
Requires-Dist: jax==0.10.1
Requires-Dist: numpy
Requires-Dist: pydantic
Requires-Dist: scipy
Provides-Extra: all
Requires-Dist: flashbax==0.1.3; extra == 'all'
Requires-Dist: flax; extra == 'all'
Requires-Dist: hydra-core==1.3.7; extra == 'all'
Requires-Dist: ipykernel; extra == 'all'
Requires-Dist: matplotlib; extra == 'all'
Requires-Dist: nbconvert; extra == 'all'
Requires-Dist: omegaconf==2.3.0; extra == 'all'
Requires-Dist: optax; extra == 'all'
Requires-Dist: orbax-checkpoint; extra == 'all'
Requires-Dist: pandas; extra == 'all'
Provides-Extra: cuda12
Requires-Dist: jax[cuda12]==0.10.1; extra == 'cuda12'
Provides-Extra: cuda13
Requires-Dist: jax[cuda13]==0.10.1; extra == 'cuda13'
Provides-Extra: dev
Requires-Dist: hydra-core==1.3.7; extra == 'dev'
Requires-Dist: omegaconf==2.3.0; extra == 'dev'
Requires-Dist: pre-commit; extra == 'dev'
Requires-Dist: pyright; extra == 'dev'
Requires-Dist: pytest; extra == 'dev'
Requires-Dist: ruff; extra == 'dev'
Provides-Extra: hydra
Requires-Dist: hydra-core==1.3.7; extra == 'hydra'
Requires-Dist: omegaconf==2.3.0; extra == 'hydra'
Provides-Extra: notebook
Requires-Dist: ipykernel; extra == 'notebook'
Requires-Dist: nbconvert; extra == 'notebook'
Requires-Dist: pandas; extra == 'notebook'
Provides-Extra: rocm
Requires-Dist: jax-rocm7-plugin==0.10.1; (sys_platform == 'linux' and platform_machine == 'x86_64') and extra == 'rocm'
Requires-Dist: jax[rocm7-local]==0.10.1; (sys_platform == 'linux' and platform_machine == 'x86_64') and extra == 'rocm'
Provides-Extra: tpu
Requires-Dist: jax[tpu]==0.10.1; (sys_platform == 'linux' and platform_machine == 'x86_64') and extra == 'tpu'
Provides-Extra: training
Requires-Dist: flashbax==0.1.3; extra == 'training'
Requires-Dist: flax; extra == 'training'
Requires-Dist: matplotlib; extra == 'training'
Requires-Dist: optax; extra == 'training'
Requires-Dist: orbax-checkpoint; extra == 'training'
Provides-Extra: viz
Requires-Dist: matplotlib; extra == 'viz'
Provides-Extra: wandb
Requires-Dist: wandb; extra == 'wandb'
Description-Content-Type: text/markdown

# MARL-BattleGrounds

**The JAX-native Benchmark for Heterogeneous and Competitive Multi-agent Reinforcement Learning.**

[![PyPI](https://img.shields.io/pypi/v/marl-battlegrounds)](https://pypi.org/project/marl-battlegrounds/)
[![Python](https://img.shields.io/pypi/pyversions/marl-battlegrounds)](https://pypi.org/project/marl-battlegrounds/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/LICENSE)
[![Website](https://img.shields.io/badge/website-marl--battlegrounds.com-2b6cb0)](https://marl-battlegrounds.com)

![Two teams of five fight on a MARL-BattleGrounds map](https://marl-battlegrounds.com/assets/art/key-art-preview.jpg)

Two teams of up to five agents fight Team Deathmatch on 52 maps. Each agent
is a Mage, Warrior, Hunter, Rogue or Priest, so a team wins by combining
movement, attacks, healing and support. The whole game is written in JAX, so
thousands of games run at once on one GPU. You bring the method; MARL-BGs
gives you the game, training helpers, evaluation, tournaments and replays.

## Install

On Linux, an Apple Silicon Mac or Windows with WSL2, one command installs everything and opens the BattleClient. It picks the right JAX build for your GPU:

```bash
curl -L https://github.com/oceansystemslab/MARL-BattleGrounds/archive/refs/tags/v1.0.0.tar.gz | tar xz
cd MARL-BattleGrounds-1.0.0 && sh start.sh
```

In your own Python project (Python 3.12 to 3.14; 3.14 recommended):

```bash
pip install "marl-battlegrounds[all]"           # CPU
pip install "marl-battlegrounds[all,cuda13]"    # NVIDIA GPU, driver 580 or newer
```

With Docker, on any system (add `--gpus all` and the `1.0.0-cuda12` tag for an NVIDIA GPU):

```bash
docker run --rm -it -p 127.0.0.1:8765:8766 -p 127.0.0.1:8767:8768 -v "${PWD}:/work" ghcr.io/oceansystemslab/marl-battlegrounds:1.0.0
```

| Route | Works on | Tested |
|---|---|---|
| One command, `sh start.sh` | Linux, Apple Silicon Mac, Windows with WSL2 | Linux with NVIDIA; install and evaluation on Linux and Mac CPUs, by every release; WSL2 not tested yet |
| pip | Linux, Apple Silicon Mac, WSL2 | Linux and Apple Silicon Mac, by every release |
| Docker | Any system with Docker; NVIDIA on Linux and Windows hosts | Linux, CPU and NVIDIA; the arm64 image by every release, under emulation; Docker Desktop not tested yet |
| Colab and Kaggle, [Open in Colab](https://colab.research.google.com/github/oceansystemslab/MARL-BattleGrounds/blob/v1.0.0/examples/quickstart.ipynb) | A browser | Not tested yet |

Docker runs everything: play, replays, training, evaluation, tournaments, analysis and notebooks. Its limits:

| Limit | What to do |
|---|---|
| No GPU in Docker on a Mac | Use the CPU image; Docker cannot reach Apple GPUs |
| `study` (big multi-run jobs) is not tested in containers and stops when the container stops | Start the container with `-d`; not tested yet |
| You cannot add packages to a running container | Build a two-line Dockerfile: `FROM ghcr.io/oceansystemslab/marl-battlegrounds:1.0.0`, then `RUN uv pip install <package>` |
| Docker Desktop (Mac, Windows) limits the container's memory | Raise the memory limit in Docker Desktop's settings for large training runs |

Every route, the GPU builds, Windows, clusters and offline use: [the install guide](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/install.md).

## Quick Start

```python
import jax
import marl_battlegrounds as marl_bgs

env = marl_bgs.make("tdm", map_id=0, num_envs=4)  # four games at once
key = jax.random.key(0)
observation, state = env.reset(key)
for _ in range(50):
    key, act_key, step_key, reset_key = jax.random.split(key, 4)
    actions = env.sample_actions(act_key, state)  # legal random actions
    observation, state, reward, done, info = env.step(step_key, state, actions)
    observation, state = env.reset_done(reset_key, state)  # restart finished games
```

On one RTX 5090 with 1,024 parallel 5v5 games, the simulator alone runs 50,416 game steps per second (random legal actions, compilation excluded), and recurrent MAPPO trains at 29,878 game steps per second in pure self-play.

## Train

One command trains recurrent IPPO with the Season 0 learner settings (250 million game
steps, a few hours on one RTX 5090) against scripted opponents, and keeps its best
checkpoint on the validation maps:

```bash
python -m marl_battlegrounds.experiments.train method=ippo pool=scripted \
  'validation_opponents=[tdm-alpha,tdm-beta,tdm-gamma]' output_dir=runs/my-ippo
```

## Evaluate

```python
import marl_battlegrounds as marl_bgs

results = marl_bgs.evaluate("random", "tdm-alpha", num_episodes=32, num_envs=32)
print(results.head_to_head())  # games, wins, draws, losses, point margin
```

Your method is always Team A, and paired games swap the spawn ends, so a map's
layout cannot favor either side.

## Watch And Play

```bash
python -m marl_battlegrounds replay path/to/game.marlbg-replay.json
```

```python
import marl_battlegrounds as marl_bgs

marl_bgs.battle_client(team_b="tdm-alpha")  # play in your browser against a scripted team
```

## The Six Built-In Learners

| Method | Name | Memory | Critic |
|---|---|---|---|
| Recurrent MAPPO | `mappo` | Recurrent | Centralized |
| Feedforward MAPPO | `ff_mappo` | None | Centralized |
| Recurrent IPPO | `ippo` | Recurrent | One per agent |
| Feedforward IPPO | `ff_ippo` | None | One per agent |
| QMIX | `qmix` | Recurrent | QMIX mixer |
| PQN-VDN | `pqn_vdn` | Recurrent | Sum of agent values |

The six trained Season 0 models, with every save from 0 to 250M steps, load by name,
for example `marl_bgs.load_method("MARL-BattleGrounds/tdm-season-0-qmix@final")`
([Training](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/training.md#use-a-released-season-0-model)).

Your own method can be anything that turns permitted observations into legal
actions: a network, a planner, a language model or a mix.

## Guides

| Guide | What It Covers | Website |
|---|---|---|
| [Getting Started](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/getting_started.md) | Install, your first games, maps, rosters, what agents see | [start](https://marl-battlegrounds.com/#start) |
| [Game Rules](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/game_rules.md) | Classes, abilities, scoring, Red Zone, scenarios | [game](https://marl-battlegrounds.com/#game) |
| [Your Method](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/your_method.md) | Plug in your own model or learner, with JAX's jit, vmap and scan | [your-method](https://marl-battlegrounds.com/#your-method) |
| [Training](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/training.md) | The six learners, opponents, curriculum, rewards, checkpoints | [training](https://marl-battlegrounds.com/#training) |
| [Hydra Configs](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/hydra.md) | Settings files, overrides, sweeps, studies and the experiment stages | [configs](https://marl-battlegrounds.com/#configs) |
| [Evaluation](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/evaluation.md) | Fair games, validation and test maps, scenarios, saved results | [evaluate](https://marl-battlegrounds.com/#evaluate) |
| [Metrics](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/metrics.md) | What is measured and how to read it | [metrics](https://marl-battlegrounds.com/#metrics) |
| [Tournaments](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/tournaments.md) | Round robins, Elo ratings and tiers | [round-robin](https://marl-battlegrounds.com/#round-robin) |
| [Tournament Rules](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/tournament_rules.md) | The leaderboard rulebook | [rules](https://marl-battlegrounds.com/#rules) |
| [Replays](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/replays.md) | Save games and watch them in the Replay Viewer | [replay-format](https://marl-battlegrounds.com/#replay-format) |
| [BattleClient](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/battle_client.md) | Play against or beside your own agents in the browser | [battleclient](https://marl-battlegrounds.com/#battleclient) |
| [LLM Agents](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/llm.md) | Run a language model as a team | [llm](https://marl-battlegrounds.com/#llm) |
| [Performance](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/performance.md) | Speed and memory, and how to measure them on your machine | [performance](https://marl-battlegrounds.com/#performance) |
| [Method Sources](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/method_sources.md) | Where the built-in learners come from | [method-sources](https://marl-battlegrounds.com/#method-sources) |
| [API Reference](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/docs/api.md) | Every public name on one page | [api](https://marl-battlegrounds.com/#api) |

## Reproduce The Paper

The paper's commands, frozen settings and result tables are in
[`scripts/paper_experiments/`](https://github.com/oceansystemslab/MARL-BattleGrounds/tree/main/scripts/paper_experiments).

## Cite

Use the "Cite this repository" button on GitHub, or:

```bibtex
@software{marl_battlegrounds,
  author = {Hawili, Ulixes Tariq},
  title = {{MARL-BattleGrounds}: The JAX-native Benchmark for Heterogeneous and Competitive Multi-agent Reinforcement Learning},
  version = {1.0.0},
  year = {2026},
  url = {https://github.com/oceansystemslab/MARL-BattleGrounds}
}
```

## Contribute

Pull requests are welcome. Start with
[CONTRIBUTING.md](https://github.com/oceansystemslab/MARL-BattleGrounds/blob/main/CONTRIBUTING.md),
and open an [issue](https://github.com/oceansystemslab/MARL-BattleGrounds/issues)
for bugs, ideas, questions and leaderboard submissions.

## Author

MARL-BattleGrounds is developed by Ulixes Tariq Hawili as part of the SPADS CDT,
University of Edinburgh School of Engineering, and the Ocean Systems Lab at HWU.

## License

Licensed under the Apache License, Version 2.0.
