Metadata-Version: 2.4
Name: evorl-jax
Version: 0.1.0
Summary: An efficient Evolutionary Reinforcement Learning Framework on Jax
Author-email: Bowen Zheng <bowen.zheng@protonmail.com>
Project-URL: Homepage, https://github.com/EMI-Group/evorl
Project-URL: Documentation, https://evorl.readthedocs.io/
Project-URL: Repository, https://github.com/EMI-Group/evorl
Project-URL: Issues, https://github.com/EMI-Group/evorl/issues
Keywords: jax,reinforcement-learning,evolutionary-computation
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: GPU
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3.10
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: aim
Requires-Dist: brax
Requires-Dist: chex
Requires-Dist: distrax
Requires-Dist: flax
Requires-Dist: tfp-nightly[jax]
Requires-Dist: gymnasium>=1.1.0
Requires-Dist: hydra-core
Requires-Dist: hydra-joblib-launcher
Requires-Dist: jax>=0.4.35
Requires-Dist: jaxlib>=0.4.35
Requires-Dist: numpy
Requires-Dist: omegaconf
Requires-Dist: optax
Requires-Dist: orbax-checkpoint
Requires-Dist: pandas
Requires-Dist: scipy
Requires-Dist: typing-extensions
Provides-Extra: wandb
Requires-Dist: wandb; extra == "wandb"
Provides-Extra: swanlab
Requires-Dist: swanlab; extra == "swanlab"
Provides-Extra: comet
Requires-Dist: comet_ml; extra == "comet"
Provides-Extra: neptune
Requires-Dist: neptune-scale; extra == "neptune"
Provides-Extra: gymnax
Requires-Dist: gymnax; extra == "gymnax"
Provides-Extra: jumanji
Requires-Dist: jumanji; extra == "jumanji"
Provides-Extra: jaxmarl
Requires-Dist: jaxmarl; extra == "jaxmarl"
Provides-Extra: mujoco-playground
Requires-Dist: playground; extra == "mujoco-playground"
Provides-Extra: envpool
Requires-Dist: envpool; extra == "envpool"
Provides-Extra: gymnasium
Requires-Dist: gymnasium[atari,box2d,classic-control,mujoco]>=1.1.0; extra == "gymnasium"
Provides-Extra: dev
Requires-Dist: furo; extra == "dev"
Requires-Dist: myst-parser; extra == "dev"
Requires-Dist: pre-commit; extra == "dev"
Requires-Dist: pytest; extra == "dev"
Requires-Dist: ruff; extra == "dev"
Requires-Dist: sphinx; extra == "dev"
Requires-Dist: sphinx-autodoc2; extra == "dev"
Requires-Dist: sphinx-copybutton; extra == "dev"
Dynamic: license-file

<h1 align="center">
  <a href="https://github.com/EMI-Group/evox">
    <picture>
      <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/evox_logo_dark.svg">
      <source media="(prefers-color-scheme: light)" srcset="https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/evox_logo_light.svg">
      <img alt="EvoX Logo" height="50" src="https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/evox_logo_light.svg">
    </picture>
  </a>
</h1>

<p align="center">
  <img src="https://github.com/google/brax/raw/main/docs/img/humanoid_v2.gif", width=160, height=160/>
  <img src="https://github.com/kenjyoung/MinAtar/raw/master/img/breakout.gif", width=160, height=160>
  <img src="https://raw.githubusercontent.com/instadeepai/jumanji/main/docs/env_anim/bin_pack.gif", width=160, height=160>
</p>

<h2 align="center">
  <p>🌟 EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning 🌟</p>
  <a href="https://arxiv.org/abs/2501.15129">
    <img src="https://img.shields.io/badge/paper-arxiv-red?style=for-the-badge" alt="EvoRL Paper on arXiv">
  </a>
</h2>


# Table of Contents
- [Table of Contents](#table-of-contents)
- [Introduction](#introduction)
  - [Highlight](#highlight)
    - [Update](#update)
  - [Documentation](#documentation)
  - [Overview of Key Concepts in EvoRL](#overview-of-key-concepts-in-evorl)
- [Installation](#installation)
- [Quickstart](#quickstart)
  - [Training](#training)
  - [Logging](#logging)
  - [Env Rendering](#env-rendering)
- [Algorithms](#algorithms)
- [RL Environments](#rl-environments)
  - [Current Supported Environments](#current-supported-environments)
- [Performance](#performance)
- [Bug report \& Discussion](#bug-report--discussion)
- [Acknowledgement](#acknowledgement)
  - [Citing EvoRL](#citing-evorl)


# Introduction

EvoRL is a fully GPU-accelerated framework for Evolutionary Reinforcement Learning, which is implemented by JAX and provides end-to-end GPU-accelerated training pipelines, including following processes:

- Reinforcement Learning (RL)
- Evolutionary Computation (EC)
- Environment Simulation

EvoRL provides a highly efficient and user-friendly platform to develop and evaluate RL, EC and EvoRL algorithms.

> [!NOTE]
> EvoRL is a sister project of [EvoX](https://github.com/EMI-Group/evox).

## Highlight

- **End-to-end training pipelines**: The training pipelines for RL, EC and EvoRL are entirely executed on GPUs, eliminating dense communication between CPUs and GPUs in traditional implementations and fully utilizing the parallel computing capabilities of modern GPU architectures.
  - Most algorithms has a `Workflow.step()` function that is capable of `jax.jit` and `jax.vmap()`, supporting parallel training and JIT on full computation graph.
- **Easy integration between EC and RL**: Due to modular design, EC components can be easily plug-and-play in workflows and cooperate with RL.
- **Implementation of EvoRL algorithms**: Currently, we provide two popular paradigms in Evolutionary Reinforcement Learning: Evolution-guided Reinforcement Learning (ERL): ERL, CEM-RL; and Population-based AutoRL: PBT.
- **Unified Environment API**: Support multiple GPU-accelerated RL environment packages (eg: Brax, gymnax, ...). Multiple Env Wrappers are also provided.
- **Object-oriented functional programming model**: Classes define the static execution logic and their running states are stored externally.

### Update

- 2025-07-14: Our paper *"EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning"* is accepted by ACM TELO.

- 2025-04-01: Add support for Mujoco Playground Environments.

## Documentation

- For comprehensive guidance, please visit our [Documentation](https://evorl.readthedocs.io/latest/), where you'll find detailed installation steps, tutorials, practical examples, and complete API references.

- EvoRL is also indexed by DeepWiki, providing an AI assistant for beginners. Feel free to ask any question about this repo at https://deepwiki.com/EMI-Group/evorl.

## Overview of Key Concepts in EvoRL

![](https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/evorl_arch.svg)

- **Workflow** defines the training logic of algorithms.
- **Agent** defines the behavior of a learning agent, and its optional loss functions.
- **Env** provides a unified interface for different environments.
- **SampleBatch** is a data structure for continuous trajectories or shuffled transition batch.
- **EC** module provide EC components like Evolutionary Algorithms (EAs) and related operators.



# Installation

EvoRL is developed on the top of `jax`. So `jax` should be installed first, please follow [JAX official installation guide](https://jax.readthedocs.io/en/latest/quickstart.html#installation). Install the released package from PyPI (available after the first release):

```shell
pip install evorl-jax
```

The distribution name is `evorl-jax`; the Python import remains `import evorl`.
For the latest development version and the training scripts/configs used below, install from source:

```shell
# Install the evorl package from source
git clone https://github.com/EMI-Group/evorl.git
cd evorl
pip install -e .
```

Aim is included for default experiment logging. WandB, SwanLab, Comet, and Neptune
are optional; install their SDKs directly or use EvoRL extras. See
[Experiment Logging installation](https://evorl.readthedocs.io/latest/guide/installation.html#experiment-logging).

For developers, see [Contributing to EvoRL](https://evorl.readthedocs.io/latest/dev/contributing.html)

# Quickstart

## Training

EvoRL uses [hydra](https://hydra.cc/) to manage configs and run algorithms. Users can use `scripts/train.py` or `scripts/train_dist.py` to run algorithms from CLI.

```text
# hierarchy of folder `configs/`
configs
├── agent
│   ├── ppo.yaml
│   ├── ...
...
├── config.yaml
├── env
│   ├── brax
│   │   ├── ant.yaml
│   │   ├── ...
│   ├── envpool
│   └── gymnax
└── logging.yaml
```

Specify the `agent` and `env` field based on the related config file path (`*.yaml`) in `configs` folder. For example: To train the *PPO* agent with config file in `configs/agent/ppo.yaml` on the Brax environment *Ant* with config file in `configs/env/brax/ant.yaml`, use:

```shell
python scripts/train.py agent=ppo env=brax/ant

# Parallel training two seeds on each GPU.
CUDA_VISIBLE_DEVICES=0,5 python scripts/train_dist.py -m hydra/launcher=joblib \
    agent=exp/ppo/brax/ant env=brax/ant seed=114,514
```

If multiple GPUs are detected, most algorithms will be automatically trained in distributed mode.

For more advanced usage, see our documentation: [Training](https://evorl.readthedocs.io/latest/guide/quickstart.html#advanced-usage).

## Logging

With the default Hydra configuration, a single run stores outputs in `outputs/<script>/<timestamp>/`, and [multi-run mode](https://hydra.cc/docs/tutorials/basic/running_your_app/multi-run/) (`-m`) uses `multirun/<script>/<timestamp>/<overrides>/`, where `<script>` is `train` or `train_dist`. `LogRecorder` writes `<experiment-name>.log` there; checkpoints use the `checkpoints/` subdirectory when `checkpoint.enable=true`.

By default, the training scripts enable `LogRecorder` and `AimRecorder` (`recorders: [log, aim]`). Aim stores runs locally in the shared `aim/.aim` repository under the directory where training was launched. View and compare runs from that directory:

```shell
aim up --repo aim
```

To use WandB, install its optional extra (or run `pip install wandb`) and select it explicitly:

```shell
pip install -e ".[wandb]"
wandb login
python scripts/train.py agent=ppo env=brax/ant 'recorders=[log,wandb]'
```

The supported recorder names are `log`, `aim`, `wandb`, `swanlab`, `comet`, and `neptune`. Multiple installed backends can be selected together, for example `'recorders=[log,aim,wandb]'`. See [Logging](https://evorl.readthedocs.io/latest/guide/quickstart.html#logging) for installation, grouping, and backend behavior.

Example dashboard when using the optional WandB recorder:

![](https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/evorl_wandb.png)

## Env Rendering

We provide some example visualization scripts for brax and playground environments: [visualize_mjx.ipynb](./scripts/visualize_mjx.ipynb).

# Algorithms

Currently, EvoRL supports 4 types of algorithms

| Type                    | Algorithms                                                                                                    |
| ----------------------- | ------------------------------------------------------------------------------------------------------------- |
| RL                      | A2C, PPO, IMPALA, DQN, DDPG, TD3, SAC, TD7                                                                    |
| EA                      | OpenES, VanillaES, ARS, CMA-ES, algorithms from [EvoX](https://github.com/EMI-Group/evox) (PSO, NSGA-II, ...) |
| Evolution-guided RL     | ERL-GA, ERL-ES, ERL-EDA, CEMRL, CEMRL-OpenES                                                                  |
| Population-based AutoRL | PBT family (e.g: PBT-PPO, PBT-SAC, PBT-CSO-PPO)                                                               |

# RL Environments

By default, `pip install evorl-jax` will automatically install environments on `brax`. If you want to use other supported environments, please install the additional environment packages. We provide useful extras for different environments.

For example:

```shell
# ===== GPU-accelerated Environments =====
# Mujoco playground Envs:
pip install -e ".[mujoco-playground]"
# gymnax Envs:
pip install -e ".[gymnax]"
# Jumanji Envs:
pip install -e ".[jumanji]"
# JaxMARL Envs:
pip install -e ".[jaxmarl]"

# ===== CPU-based Environments =====
# EnvPool Envs:
pip install -e ".[envpool]"
# Gymnasium Envs:
pip install -e ".[gymnasium]"
```

> [!WARNING]
> These additional environments have limited supports and some algorithms are incompatible with them.

## Current Supported Environments

| Environment Library                                                        | Descriptions                            |
| -------------------------------------------------------------------------- | --------------------------------------- |
| [Brax](https://github.com/google/brax)                                     | Robotic control                         |
| [MuJoCo Playground](https://github.com/google-deepmind/mujoco_playground)  | Robotic control                         |
| [gymnax (experimental)](https://github.com/RobertTLange/gymnax)            | classic control, bsuite, MinAtar        |
| [JaxMARL (experimental)](https://github.com/FLAIROx/JaxMARL)               | Multi-agent Envs                        |
| [Jumanji (experimental)](https://github.com/instadeepai/jumanji)           | Game, Combinatorial optimization        |
| [EnvPool (experimental)](https://github.com/sail-sg/envpool)               | High-performance CPU-based environments |
| [Gymnasium (experimental)](https://github.com/Farama-Foundation/Gymnasium) | Standard CPU-based environments         |

PRs for other environment libraries are welcomed.

# Performance

Test settings:

- Hardware:
  - 2x Intel Xeon Gold 6132 (56 logical cores in total)
  - 128 GiB RAM
  - 1x Nvidia RTX 3090
- Task: Swimmer

![](https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/es-perf.png)
![](https://raw.githubusercontent.com/EMI-Group/evorl/main/docs/_static/erl-pbt-perf.png)

# Bug report & Discussion

To keep our project organized, please use the appropriate GitHub section:

- [Issues](https://github.com/EMI-Group/evorl/issues) – For reporting **bugs** and **PR** only. When submitting an issue, please provide clear details to help with troubleshooting.
- [Discussions](https://github.com/EMI-Group/evorl/discussions) – For general questions, feature requests, and other topics.

Before posting, kindly check existing issues and discussions to avoid duplicates. Thank you for your contributions!

# Acknowledgement

- [acme](https://github.com/google-deepmind/acme)
- [EvoX](https://github.com/EMI-Group/evox)
- [Brax](https://github.com/google/brax)
- [MuJoCo Playground](https://github.com/google-deepmind/mujoco_playground)
- [Jumanji](https://github.com/instadeepai/jumanji)
- [JaxMARL](https://github.com/FLAIROx/JaxMARL)
- [gymnax](https://github.com/RobertTLange/gymnax)
- [EnvPool](https://github.com/sail-sg/envpool)

## Citing EvoRL

If you use EvoRL in your research and want to cite it in your work, please use:

```
@article{zheng2025evorl,
  author  = {Zheng, Bowen and Cheng, Ran and Tan, Kay Chen},
  doi     = {10.1145/3750053},
  journal = {ACM Trans. Evol. Learn. Optim.},
  month   = aug,
  title   = {EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning},
  url     = {https://doi.org/10.1145/3750053},
  year    = {2025}
}
```
