Metadata-Version: 2.4
Name: matra-genoa
Version: 0.1.0
Summary: Matra Genoa: transformer-based crystal structure generation utilities.
Author: Pierre-Paul De Breuck, Hashim A. Piracha, Gian-Marco Rignanese, Miguel A. L. Marques
License: Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to use,
        copy, modify, and distribute the Software for non-commercial research,
        education, evaluation, and personal use, subject to the following conditions:
        
        1. The above copyright notice and this permission notice shall be included in
           all copies or substantial portions of the Software.
        2. Any modified version of the Software must clearly state that changes were
           made.
        3. The Software may not be sold, sublicensed, or otherwise used for a
           commercial purpose without prior written permission from the copyright
           holder.
        4. "Commercial purpose" includes use of the Software or any derivative work in
           exchange for payment, as part of a paid service, or in support of a
           revenue-generating product or workflow.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE, AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT, OR OTHERWISE, ARISING FROM,
        OUT OF, OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/ppdebreuck/matra-genoa
Project-URL: Repository, https://github.com/ppdebreuck/matra-genoa
Project-URL: Issues, https://github.com/ppdebreuck/matra-genoa/issues
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Chemistry
Classifier: Topic :: Scientific/Engineering :: Physics
Classifier: License :: Other/Proprietary License
Classifier: Operating System :: OS Independent
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24
Requires-Dist: pandas>=2.0
Requires-Dist: pymatgen>=2024.1
Requires-Dist: scipy>=1.10
Requires-Dist: sympy>=1.12
Requires-Dist: torch>=2.1
Requires-Dist: tqdm>=4.66
Provides-Extra: dev
Requires-Dist: black>=24.0; extra == "dev"
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: ruff>=0.5; extra == "dev"
Provides-Extra: tutorials
Requires-Dist: jupyter>=1.0; extra == "tutorials"
Requires-Dist: matplotlib>=3.8; extra == "tutorials"
Provides-Extra: materials
Requires-Dist: h5py>=3.10; extra == "materials"
Requires-Dist: smact>=2.5; extra == "materials"
Dynamic: license-file

# Matra-Genoa

[![arXiv](https://img.shields.io/badge/arXiv-2501.16051-lightgrey)](https://arxiv.org/abs/2501.16051)

**Matra-Genoa** is a generative material transformer for the efficient generation of novel, symmetry-aware crystal structures. It utilizes an invertible tokenized representation of symmetrized crystals, including free coordinates. It can be conditioned on stability (energy above the convex hull), elemental compositions, space group and Wyckoff positions.

This repo contains the source code to (conditionally) generate crystal structures building on PyTorch, Pymatgen and more.

---

## Resources
- 📄 **Paper:** [arXiv:2501.16051](https://arxiv.org/abs/2501.16051)
- 📊 **MatraGenoa3M:** 3 million generated crystals on [figshare](https://doi.org/10.6084/m9.figshare.28271294.v1). Not relaxed; future updates may include novel relaxed structures.
- 🖥️ **Matra-Genoa-MPAS Front End:** [matra.pierrepauldb.com](https://matra.pierrepauldb.com)

---

## Citation

If you use this model or code in your research, please cite:

> Pierre-Paul De Breuck, Hashim A. Piracha, Gian-Marco Rignanese, Miguel A. L. Marques  
> *A generative material transformer using Wyckoff representation*  
> arXiv:2501.16051 (2025) – [https://arxiv.org/abs/2501.16051](https://arxiv.org/abs/2501.16051)

> De Breuck, P.-P., Piracha, H.A., Rignanese, G.-M. and Marques, M.A.L. A generative material transformer using Wyckoff representation. *npj Comput Mater* **12**, 60 (2026). https://doi.org/10.1038/s41524-025-01940-8

## Quickstart (Recommended)

Get started quickly by setting up a virtual environment and installing the package:

```bash
# Create and activate a virtual environment (uv recommended)
uv venv
source .venv/bin/activate

# Install the package
uv pip install matra-genoa
```

Once installed, you can start generating crystal structures immediately:

```python
from matra_genoa import MatraGenoa

# Initialize the model (downloads default checkpoints automatically)
model = MatraGenoa()

# Generate tokens for 4 structures
sequences = model.generate(n=4, T=0.75, batch_size=16)

# Reconstruct a generated sequence into a pymatgen structure
from matra_genoa.utils import reconstruct
structure = reconstruct(sequences[0])
print(structure)
```

> [!IMPORTANT]
> Generated crystal structures should be considered "as-is" from the transformer. It is highly recommended to relax these structures using a cheap universal Machine Learning Interatomic Potential (uMLIP) after reconstruction.

### GPU Acceleration

Generating crystal structures is significantly faster on a GPU. You can move the model to your preferred device:

```python
import torch

# Check for CUDA (NVIDIA) or MPS (Apple Silicon)
if torch.cuda.is_available():
    device = torch.device("cuda")
elif torch.backends.mps.is_available():
    device = torch.device("mps")
else:
    device = torch.device("cpu")

print(f"Using device: {device}")
model.to(device)
```

> [!TIP]
> When using a GPU, you can increase the `batch_size` in `model.generate(..., batch_size=256)` to improve throughput.

> [!TIP]
> **Sampling Temperature (`T`):** Use `T ≈ 0.7` for structures close to the training distribution (generally more stable). For greater diversity, you can increase this up to `T ≈ 2.0`, though higher temperatures are more likely to produce invalid or "broken" crystal structures.

### Conditioned Generation

You can guide the generation process by providing a starting prompt (conditioning). The model uses specific syntax blocks followed by a `STOP` token to define constraints.


#### Available Syntax Blocks:
- `AMT [value]`: Number of distinct chemical elements (e.g., `AMT 2` for binaries).
- `EHULL_DISC [EH0/EH1]`: Discrete stability target. `EH0` indicates stability (distance to convex hull < 75 meV/atom).
- `ELMS [elements]`: Target chemical elements.
- `STOICH [ratios]`: Target stoichiometry.
- `SPACEGROUP S[1-230]`: Target space group number.
- `WYCKOFF W[idx]`: Target Wyckoff positions.
- `EHULL [value]`: Continuous stability target.

#### Example:

```python
# Target stable binary compounds containing Na and Cl
condition = "EHULL_DISC EH0 STOP AMT 2 STOP ELMS Na Cl STOP"
tokens = model.generate(n=100, condition=condition)
```

### Batch Generation and Decoding

For generating many structures at once, `generate_structures()` samples and reconstructs sequences into `pymatgen` `Structure` objects in one call, keeping only sequences that reconstruct successfully:

```python
structures = model.generate_structures(n=1000, T=0.8, decode_jobs=8, progress=True)
```

`decode_jobs` sets how many CPU workers reconstruct structures in parallel. Decoding is always CPU-bound, so for best throughput put the model on a GPU (see [GPU Acceleration](#gpu-acceleration)) for generation while `decode_jobs` handles decoding on CPU. You can also decode a sequence list you already have with `model.decode(sequences, n_jobs=8)`. See [`03_batch_generation_and_decoding.ipynb`](tutorials/03_batch_generation_and_decoding.ipynb) for more detail.

## Tutorials

For more in-depth examples and advanced usage, explore the interactive notebooks:

- [`01_load_model_and_generate.ipynb`](tutorials/01_load_model_and_generate.ipynb): Basics of loading, sampling, and structure reconstruction.
- [`02_conditioned_generation_and_reconstruction.ipynb`](tutorials/02_conditioned_generation_and_reconstruction.ipynb): Guided generation with conditioning prompts.
- [`03_batch_generation_and_decoding.ipynb`](tutorials/03_batch_generation_and_decoding.ipynb): CPU-parallel decoding of large batches with `generate_structures()`/`decode()`.
- [`examples/quickstart.py`](examples/quickstart.py): A ready-to-run Python script.

---

## Detailed Installation

### From PyPI

```bash
pip install matra-genoa
```

### From GitHub

You can also install the package directly from the source repository:

```bash
pip install git+https://github.com/ppdebreuck/matra-genoa.git
```

To include optional dependencies for development or extra materials helpers:

```bash
# For tutorials and development
pip install "matra-genoa[dev,tutorials] @ git+https://github.com/ppdebreuck/matra-genoa.git"

# For additional materials science helpers (smact, etc.)
pip install "matra-genoa[materials] @ git+https://github.com/ppdebreuck/matra-genoa.git"
```

### Local Development

If you have the repository cloned locally, install it in editable mode:

```bash
pip install -e ".[dev,tutorials]"
```

*(To generate a source distribution locally for hosting, run `python -m build` which will create a `dist/` directory.)*

## Checkpoints

The following models are currently available:

- `Matra-Genoa-MPAS` (Default)
- `Matra-Genoa-MP`

Checkpoints are automatically downloaded to your local cache:
- `$MATRA_CACHE_DIR` if set
- otherwise `~/.cache/matra` (on Linux/macOS)

## Package Layout

The repository follows a modern `src/` layout:

- `src/matra_genoa/`: Core Python package logic
- `tests/`: Unit tests
- `tutorials/`: Interactive Jupyter notebooks
- `examples/`: Script-based examples
- `checkpoints/`: Local checkpoint placeholders and cache logic
