Metadata-Version: 2.4
Name: persistbench
Version: 0.1.1
Summary: Model-independent video persistence inference and evaluation
License-Expression: MIT
Project-URL: Homepage, https://www.guangzhaohe.com/persistbench/
Project-URL: Repository, https://github.com/guangzhaohe/PersistBench
Project-URL: Documentation, https://github.com/guangzhaohe/PersistBench#readme
Project-URL: Issues, https://github.com/guangzhaohe/PersistBench/issues
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy<3,>=1.24
Requires-Dist: Pillow>=9
Requires-Dist: PyYAML>=6
Requires-Dist: easydict>=1.10
Requires-Dist: loguru>=0.7
Requires-Dist: tqdm>=4.65
Provides-Extra: baselines
Requires-Dist: opencv-python-headless>=4.8; extra == "baselines"
Provides-Extra: eval
Requires-Dist: torch>=2.5; extra == "eval"
Requires-Dist: torchvision>=0.20; extra == "eval"
Requires-Dist: opencv-python-headless>=4.8; extra == "eval"
Requires-Dist: scipy>=1.13; extra == "eval"
Requires-Dist: imageio>=2.31; extra == "eval"
Requires-Dist: imageio-ffmpeg>=0.4; extra == "eval"
Requires-Dist: pyzmq>=25; extra == "eval"
Requires-Dist: rich>=13; extra == "eval"
Requires-Dist: openai>=1.55; extra == "eval"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: build>=1; extra == "dev"
Dynamic: license-file

<div align="center">
<h1 align="center"><img src="https://raw.githubusercontent.com/guangzhaohe/PersistBench/main/assets/persistbench-icon.png" alt="" width="108" height="54" align="absmiddle"> PersistBench: Can 4D Foundation Models Remember?</h1>

[![arXiv](https://img.shields.io/badge/arXiv-2609.20819-b31b1b?logo=arxiv&logoColor=white)](https://arxiv.org/abs/2609.20819)
[![PersistBench Video](https://img.shields.io/badge/PersistBench-Video-c4302b?logo=youtube&logoColor=red)](https://www.youtube.com/watch?v=xWpxP3dNO68)
[![PersistBench Website](https://img.shields.io/badge/PersistBench-Website-green?logo=googlechrome&logoColor=green)](https://www.guangzhaohe.com/persistbench/)
[![Leaderboard](https://img.shields.io/badge/Leaderboard-blue?logo=googleanalytics&logoColor=white)](https://www.guangzhaohe.com/persistbench/#leaderboard)
[![Dataset Code](https://img.shields.io/badge/Dataset-Code-blue?logo=github)](https://github.com/guangzhaohe/PersistBenchDataGen)
![Visitors](https://hitscounter.dev/api/hit?url=https%3A%2F%2Fwww.guangzhaohe.com%2Fpersistbench%2F&label=Visitors&icon=people&color=%23FFA500)

**NeurIPS 2026 Evaluations and Datasets (Spotlight)**

[Guangzhao He](https://www.guangzhaohe.com/), [Hadar Averbuch-Elor](https://www.hadarelor.com/)\*, [Wei-Chiu Ma](https://www.cs.cornell.edu/~weichiu/)\* ("\*" denotes equal advising)
</div>

## Overview

PersistBench is a dataset and metric suite for evaluating visual memory in 4D
foundation models, including camera-controlled video generators and 4D reconstruction
models. It uses 360° videos as reference observations to assess objects after they
leave the input camera's view, measuring **object permanence**, **motion continuity**,
and **appearance preservation**. See the [paper abstract](https://arxiv.org/abs/2609.20819).

This early code release supports your own model and the [reference diagnostic](https://github.com/guangzhaohe/PersistBench/blob/main/baselines/reference/README.md).

## Release checklist

- [x] Custom model inference and evaluation
- [ ] Dataset download workflow
- [x] PyPI publication
- [ ] Release code for baselines in the paper

## Setup

Use Python **3.10+** in your model environment. Replace `/path/to/...` and `your-model-env` with your own paths and Conda environment name.

1. **Clone PersistBench** into your workspace:
   ```bash
   git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
   ```
2. **Activate your model's existing Conda environment**, then install from the **PersistBench repo**:
   ```bash
   conda activate your-model-env
   cd /path/to/PersistBench
   python -m pip install -e .
   ```
   Alternatively, install from PyPI in your model environment: `python -m pip install persistbench`.

**PyPI only (0.1.1+):** run `persistbench docs` for the complete workflow, or
`persistbench docs --output /path/to/persistbench-workspace` to export the guides,
model example, and Qwen launcher. Follow the
[PyPI guide](https://github.com/guangzhaohe/PersistBench/blob/main/docs/pip.md) without cloning this repo.

## Dataset

Prepare the dataset separately at `/data/benchmark`. The download workflow is coming;
currently use the [dataset format](https://github.com/guangzhaohe/PersistBench/blob/main/docs/data-format.md).

## Inference

3. **Go to your model's repo**, keeping your model environment active:
   ```bash
   cd /path/to/your-model
   ```
4. **Create `my_model.py` in that repo.** Replace the loading and generation calls
   below with your model's API; set its frame count and resolution:
   ```python
   from persistbench import BaseModel, ModelOutput

   class MyModel(BaseModel):
       num_frames = 49
       resolution = (384, 512)  # height, width; -1 preserves dataset shape

       def __init__(self, device="cuda:0"):
           super().__init__(device=device)
           self.model = load_your_model(device=device)

       def predict(self, sample):
           frames = self.model.generate(sample.input_frames,
                                        sample.input_cameras, sample.target_cameras)
           return ModelOutput(frames)  # CPU NumPy RGB [T, H, W, 3]
   ```
   See the [executable example](https://github.com/guangzhaohe/PersistBench/blob/main/examples/example_model.py) and [interface details](https://github.com/guangzhaohe/PersistBench/blob/main/docs/usage.md).
5. **Run inference from your model's repo**, in the same model environment:
   ```bash
   persistbench infer --model ./my_model.py:MyModel --name my-model --dataset /data/benchmark
   ```
   Predictions are saved to `/path/to/your-model/outputs/null/my-model/`.

## Evaluation

1. **Create the evaluation Conda environment.** Run from the **PersistBench repo**:
   ```bash
   cd /path/to/PersistBench
   conda create -n persistbench-eval python=3.11 pip -y
   conda activate persistbench-eval
   python -m pip install '.[eval]'
   ```
2. **Install the metric models and download their checkpoints**, in this same environment
   and repo: follow the complete [checkpoint guide](https://github.com/guangzhaohe/PersistBench/blob/main/docs/checkpoints.md) (SAM2, DINOv3, DINOv2).
3. **Start or connect to Qwen.** Follow the independent [Qwen judge guide](https://github.com/guangzhaohe/PersistBench/blob/main/serving/qwen/README.md).
   Hosting uses a separate Conda environment and terminal; keep the evaluation terminal open.
4. **Evaluate and aggregate.** In the **evaluation terminal**, from the **PersistBench repo**:
   ```bash
   cd /path/to/PersistBench
   conda activate persistbench-eval
   export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt"
   export DINOV3_REPO="$PWD/vendor/dinov3"
   export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth"
   export PERSISTBENCH_VLM_BASE_URL=http://127.0.0.1:8000/v1
   persistbench evaluate --dataset /data/benchmark --predictions /path/to/your-model/outputs/null/my-model
   persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
   ```
   For a remote Qwen server, replace the endpoint above with its URL.

Read `/path/to/your-model/outputs/persist_bench/my-model/summary.json` for
**static/dynamic × visible/invisible** averages and valid-case counts.
Evaluation saves final scores; aggregation only averages them.
See [score definitions](https://github.com/guangzhaohe/PersistBench/blob/main/docs/metrics.md) for formulas and missing-score handling.
Commands use the active Conda environment; they do not switch environments.

## Visualizations and progress

Per-case evaluation folders include matching images, tracked-mask overlays, object crops,
and judge inputs. From the **PersistBench repo**, in the **evaluation environment**:

```bash
persistbench dashboard --output /path/to/your-model/outputs --port 8080
```

Open `http://127.0.0.1:8080`. For splitting, resuming, and configuration, see
[usage](https://github.com/guangzhaohe/PersistBench/blob/main/docs/usage.md); for checks and limitations, see [validation](https://github.com/guangzhaohe/PersistBench/blob/main/docs/validation.md).

## Acknowledgements

We thank the authors of [SAM2](https://github.com/facebookresearch/sam2),
[DINOv3](https://github.com/facebookresearch/dinov3),
[DINOv2](https://github.com/facebookresearch/dinov2), and
[Qwen](https://github.com/QwenLM/Qwen3.8) for their open-source projects.

## License

This project is licensed under the [MIT License](https://github.com/guangzhaohe/PersistBench/blob/main/LICENSE).

## Citation

If you use PersistBench in your research, please cite:

```bibtex
@misc{he2026persistbench,
  title={Can {4D} Foundation Models Remember?},
  author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
  year={2026},
  eprint={2609.20819},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.20819}
}
```
