Metadata-Version: 2.4
Name: persistbench
Version: 0.1.6
Summary: Model-independent video persistence inference and evaluation
License-Expression: MIT
Project-URL: Homepage, https://www.guangzhaohe.com/persistbench/
Project-URL: Repository, https://github.com/guangzhaohe/PersistBench
Project-URL: Documentation, https://github.com/guangzhaohe/PersistBench#readme
Project-URL: Issues, https://github.com/guangzhaohe/PersistBench/issues
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy
Requires-Dist: Pillow
Requires-Dist: PyYAML
Requires-Dist: easydict
Requires-Dist: loguru
Requires-Dist: tqdm
Provides-Extra: data
Requires-Dist: numpy>=2; extra == "data"
Requires-Dist: pyarrow>=18; extra == "data"
Requires-Dist: opencv-python-headless>=4.10; extra == "data"
Requires-Dist: rich>=13; extra == "data"
Requires-Dist: yt-dlp[default]>=2025.11.12; extra == "data"
Requires-Dist: imageio-ffmpeg>=0.5; extra == "data"
Requires-Dist: huggingface_hub>=0.32; extra == "data"
Provides-Extra: eval
Requires-Dist: torch>=2.5; extra == "eval"
Requires-Dist: torchvision>=0.20; extra == "eval"
Requires-Dist: opencv-python-headless>=4.8; extra == "eval"
Requires-Dist: scipy>=1.13; extra == "eval"
Requires-Dist: imageio>=2.31; extra == "eval"
Requires-Dist: imageio-ffmpeg>=0.4; extra == "eval"
Requires-Dist: pyzmq>=25; extra == "eval"
Requires-Dist: rich>=13; extra == "eval"
Requires-Dist: openai>=1.55; extra == "eval"
Requires-Dist: ftfy; extra == "eval"
Requires-Dist: omegaconf; extra == "eval"
Requires-Dist: regex; extra == "eval"
Requires-Dist: scikit-learn; extra == "eval"
Requires-Dist: submitit; extra == "eval"
Requires-Dist: termcolor; extra == "eval"
Requires-Dist: torchmetrics; extra == "eval"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: build>=1; extra == "dev"
Dynamic: license-file

<div align="center">
<h1 align="center"><img src="assets/persistbench-icon.png" alt="" width="108" height="54" align="absmiddle"> PersistBench: Can 4D Foundation Models Remember?</h1>

[![arXiv](https://img.shields.io/badge/arXiv-2609.20819-b31b1b?logo=arxiv&logoColor=white)](https://arxiv.org/abs/2609.20819)
[![PersistBench Video](https://img.shields.io/badge/PersistBench-Video-c4302b?logo=youtube&logoColor=red)](https://www.youtube.com/watch?v=xWpxP3dNO68)
[![PersistBench Website](https://img.shields.io/badge/PersistBench-Website-green?logo=googlechrome&logoColor=green)](https://www.guangzhaohe.com/persistbench/)
[![Hugging Face Leaderboard](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Leaderboard-FFD21E)](https://huggingface.co/spaces/therealguangzhaohe/PersitBenchLeaderboard)
[![Leaderboard Submission](https://img.shields.io/badge/%F0%9F%A4%97%20Leaderboard-Submission-orange)](https://therealguangzhaohe-persitbenchleaderboard.hf.space/submit/)
[![PyPI](https://img.shields.io/pypi/v/persistbench?logo=pypi&logoColor=white)](https://pypi.org/project/persistbench/)
[![Hugging Face Dataset](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-FFD21E)](https://huggingface.co/datasets/therealguangzhaohe/PersistBenchDataset)
[![Dataset Code](https://img.shields.io/badge/Dataset-Code-blue?logo=github)](https://github.com/guangzhaohe/PersistBenchDataGen)
![Visitors](https://hitscounter.dev/api/hit?url=https%3A%2F%2Fwww.guangzhaohe.com%2Fpersistbench%2F&label=Visitors&icon=people&color=%23FFA500)

**NeurIPS 2026 Evaluations and Datasets (Spotlight)**

[Guangzhao He](https://www.guangzhaohe.com/), [Hadar Averbuch-Elor](https://www.hadarelor.com/)\*, [Wei-Chiu Ma](https://www.cs.cornell.edu/~weichiu/)\* ("\*" denotes equal advising)
</div>

## Overview

PersistBench is a dataset and evaluation suite for benchmarking visual memory in 4D
foundation models, including camera-controlled video generators, video-to-360° video generators, and 4D reconstruction models. It uses 360° videos as reference observations to evaluation objects after they leave the input camera's view, measuring **object permanence**, **motion continuity**, and **appearance preservation**.

## Release checklist

- [x] Custom model inference and evaluation
- [x] Dataset download workflow
- [x] PyPI publication
- [ ] Release code for baselines in the paper

## Setup

Inference supports Python **3.9+** in your own model environment. There are two options to install the evaluation suite. After installation, both routes share the same workflow.

### Option 1: Install from PyPI

```bash
conda activate your-model-env  # replace with your model's inference environment
python -m pip install 'persistbench>=0.1.4'
persistbench docs --output /path/to/PersistBench  # save the documents to this directory
```

The export creates a new workspace containing these guides, the model example,
and the Qwen launcher.

### Option 2: Clone and install

```bash
git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
conda activate your-model-env  # replace with your model's inference environment
cd /path/to/PersistBench
python -m pip install -e .
```

## Dataset

We host our evaluation dataset metadata and reconstruction script on [Hugging Face](https://huggingface.co/datasets/therealguangzhaohe/PersistBenchDataset).

Create a separate dataset environment, from your **PersistBench workspace**:

```bash
cd /path/to/PersistBench
conda create -n persistbench-data python=3.12 pip -y
conda activate persistbench-data
conda install -c conda-forge 'nodejs>=22' -y
```

Install the dataset dependencies using your chosen route:

| Clone | PyPI |
| :---: | :---: |
| `python -m pip install -e '.[data]'` | `python -m pip install 'persistbench[data]>=0.1.4'` |

Choose one download option, in the **dataset environment**, from your **PersistBench workspace**:

### Option 1: Local YouTube download (yt-dlp)

```bash
persistbench dataset download --output ./data/persistbench_data
```

### Option 2: Google Cloud download

Install the [Google Cloud CLI](https://docs.cloud.google.com/sdk/docs/install-sdk),
then run and follow the Google sign-in and project-selection prompts:

```bash
persistbench dataset download --output ./data/persistbench_data --backend cloud
```

Source videos download on a Cloud VM; reconstruction runs locally.
See [Cloud setup](docs/dataset.md#option-2-google-cloud-download) for requirements and VM management.

Both options save data to `./data/persistbench_data` and register its
absolute path for all your Conda environments. Inference and evaluation require
registered data. See [dataset setup](docs/dataset.md) for resuming and small test runs.

## Inference

1. **Switch back to your model environment and go to your model's repo**:
   ```bash
   conda activate your-model-env
   cd /path/to/your-model
   ```
2. **Create `my_model.py` in that repo.** Replace the loading and generation calls
   below with your model's API; set its frame count and resolution:
   ```python
   from persistbench import BaseModel, ModelOutput

   class MyModel(BaseModel):
       num_frames = 49
       resolution = (384, 512)  # height, width; -1 preserves dataset shape

       def __init__(self, device="cuda:0"):
           super().__init__(device=device)
           self.model = load_your_model(device=device)

       def predict(self, sample):
           frames = self.model.generate(sample.input_frames,
                                        sample.input_cameras, sample.target_cameras)
           return ModelOutput(frames)  # CPU NumPy RGB [T, H, W, 3]
   ```
   See the [executable example](https://github.com/guangzhaohe/PersistBench/blob/main/examples/example_model.py) and [interface details](https://github.com/guangzhaohe/PersistBench/blob/main/docs/usage.md).
3. **Run inference from your model's repo**, in the same model environment:
   ```bash
   persistbench infer --model ./my_model.py:MyModel --name my-model
   ```
   Predictions are saved to `/path/to/your-model/outputs/null/my-model/`.

## Evaluation

1. **Create the evaluation Conda environment.** Run from your **PersistBench workspace**:
   ```bash
   cd /path/to/PersistBench
   conda create -n persistbench-eval python=3.11 pip -y
   conda activate persistbench-eval
   ```
   Install the evaluation dependencies using your chosen route:

   | Clone | PyPI |
   | --- | --- |
   | `python -m pip install '.[eval]'` | `python -m pip install 'persistbench[eval]>=0.1.4'` |

2. **Install the metric models and download their checkpoints**, in this same environment
   and workspace: follow the complete [checkpoint guide](https://github.com/guangzhaohe/PersistBench/blob/main/docs/checkpoints.md) (SAM2, DINOv3, DINOv2).
3. **Configure Qwen using either option** in the [Qwen judge guide](https://github.com/guangzhaohe/PersistBench/blob/main/serving/qwen/README.md):
   **Alibaba Cloud / DashScope API** with your own key (`qwen3.8-27b`), or
   **self-hosting** the pinned BF16 checkpoint (development setup: **2 × 48 GB A6000**;
   one 48 GB GPU is insufficient). Self-hosting uses a separate Conda environment and terminal.
4. **Evaluate and aggregate.** In the **evaluation terminal**, from your **PersistBench workspace**:
   ```bash
   cd /path/to/PersistBench
   conda activate persistbench-eval
   export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt"
   export DINOV3_REPO="$PWD/vendor/dinov3"
   export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth"
   persistbench evaluate --predictions /path/to/your-model/outputs/null/my-model
   persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
   ```
   Keep the Qwen variables configured in step 3 in this terminal.

Read `/path/to/your-model/outputs/persist_bench/my-model/summary.json` for
**static/dynamic × visible/invisible** averages and valid-case counts.
Evaluation saves final scores; aggregation only averages them.
See [score definitions](https://github.com/guangzhaohe/PersistBench/blob/main/docs/metrics.md) for formulas and missing-score handling.
Commands use the active Conda environment; they do not switch environments.

## Leaderboard submission

Follow the [submission guide](docs/submissions.md) to package the evaluation report from all 2,000 cases and upload it to the leaderboard. Packaging uses the files already produced by evaluation and aggregation:

```bash
persistbench package \
  --results /path/to/your-model/outputs/persist_bench/my-model \
  --output submission.zip
```

This command is included in the PyPI package; no separate helper script is needed.

## Visualizations and progress

Per-case evaluation folders include matching images, tracked-mask overlays, object crops,
and judge inputs. From your **PersistBench workspace**, in the **evaluation environment**:

```bash
persistbench dashboard --output /path/to/your-model/outputs --port 8080
```

Open `http://127.0.0.1:8080`. For splitting, resuming, and configuration, see
[usage](https://github.com/guangzhaohe/PersistBench/blob/main/docs/usage.md); for checks and limitations, see [validation](https://github.com/guangzhaohe/PersistBench/blob/main/docs/validation.md).

## Acknowledgements

We thank the authors of [SAM2](https://github.com/facebookresearch/sam2),
[DINOv3](https://github.com/facebookresearch/dinov3),
[DINOv2](https://github.com/facebookresearch/dinov2), and
[Qwen](https://github.com/QwenLM/Qwen3.8) for their open-source projects.

## License

This project is licensed under the [MIT License](https://github.com/guangzhaohe/PersistBench/blob/main/LICENSE).

Maintainers: [automatic PyPI releases](docs/publishing.md).

## Citation

If you use PersistBench in your research, please cite:

```bibtex
@misc{he2026persistbench,
  title={Can {4D} Foundation Models Remember?},
  author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
  year={2026},
  eprint={2609.20819},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.20819}
}
```
