Metadata-Version: 2.5
Name: metal-gauss
Version: 0.2.1
Summary: Differentiable 3D Gaussian Splatting rasterizer, Metal-native. No CUDA anywhere.
Project-URL: Homepage, https://github.com/nandometzger/metal-gauss
Project-URL: Source, https://github.com/nandometzger/metal-gauss
Project-URL: Measured rejections, https://github.com/nandometzger/metal-gauss/blob/main/bench/results/NEGATIVE_RESULTS.md
License: MIT
License-File: LICENSE
Keywords: 3dgs,apple-silicon,gaussian-splatting,metal,mps,radiance-fields
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Multimedia :: Graphics :: 3D Rendering
Classifier: Topic :: Scientific/Engineering :: Image Processing
Requires-Python: >=3.10
Requires-Dist: ninja>=1.11
Requires-Dist: numpy>=1.26
Requires-Dist: pillow>=10.0
Requires-Dist: plyfile>=1.0
Requires-Dist: scipy>=1.11
Requires-Dist: torch>=2.5
Provides-Extra: bench
Requires-Dist: pycolmap>=3.10; extra == 'bench'
Provides-Extra: test
Requires-Dist: pytest>=8.0; extra == 'test'
Provides-Extra: train
Requires-Dist: pycolmap>=3.10; extra == 'train'
Provides-Extra: viewer
Requires-Dist: viser<2,>=1.1; extra == 'viewer'
Description-Content-Type: text/markdown

<p align="center"><img src="https://raw.githubusercontent.com/nandometzger/metal-gauss/main/assets/logo.png" width="320"></p>

<p align="center"><b>3D Gaussian Splatting that trains on Apple Silicon. Metal kernels, no CUDA, no Xcode.</b></p>

<p align="center">
  <img alt="Apple Silicon" src="https://img.shields.io/badge/platform-Apple%20Silicon-000000?logo=apple&logoColor=white">
  <img alt="Metal" src="https://img.shields.io/badge/backend-Metal-A855F7">
  <img alt="PyTorch MPS" src="https://img.shields.io/badge/PyTorch-MPS-EE4C2C?logo=pytorch&logoColor=white">
  <img alt="Python 3.10+" src="https://img.shields.io/badge/python-3.10%2B-3776AB?logo=python&logoColor=white">
  <img alt="License MIT" src="https://img.shields.io/badge/license-MIT-22C55E">
  <img alt="No CUDA" src="https://img.shields.io/badge/CUDA-not%20required-6E7681">
  <img alt="No Xcode" src="https://img.shields.io/badge/Xcode-not%20required-6E7681">
</p>

---

3D Gaussian Splatting that **trains** on Apple Silicon. Metal compute kernels compiled at runtime,
so there is no CUDA and no Xcode — Command Line Tools are enough.

## 🏆 Quality per minute

Given N minutes, what is the best reconstruction you can get? All 8 NeRF-synthetic scenes, identical
seed cloud per scene, one evaluator on the official 200-view test split, strictly sequential.

<!-- BEGIN:budget -->
| you have | metal-gauss | msplat | Brush | spirula |
|---|---:|---:|---:|---:|
| 30 s | **21.6** | 19.9 | 12.8 | 13.8 |
| 1 min | **24.7** | 20.8 | 14.4 | 15.1 |
| 3 min | **29.8** | 22.1 | 18.4 | 17.4 |
| 6 min | **31.3** | 22.3 | 23.4 | 19.6 |
| 15 min | **31.9** | 22.4 | 26.9 | 22.1 |
| 30 min | **31.9** | 22.4 | 26.9 | 28.5 |

*Best 8-scene-mean PSNR reachable without exceeding each budget; msplat takes its better variant at each point. Em dash means the implementation produces nothing within that budget on all 8 scenes. Full per-rung ladder in [docs/BENCHMARKS.md](https://github.com/nandometzger/metal-gauss/blob/main/docs/BENCHMARKS.md).*
<!-- END:budget -->

![PSNR vs wall-clock, 8-scene mean](https://raw.githubusercontent.com/nandometzger/metal-gauss/main/bench/results/pareto_8scene.svg)

Lines trace best-achievable-by-budget. Hollow dots are measured but beaten by a cheaper run of the
same implementation. Whiskers are ±1 s.e.m. across the 8 scenes.

![Paired per-scene margin with 95% confidence intervals](https://raw.githubusercontent.com/nandometzger/metal-gauss/main/bench/results/margin_forest.svg)

Paired per scene, because every implementation ran the same 8 scenes. One interval crosses zero:
our quality margin over spirula at 15 k is not resolved by 8 scenes, though the wall-clock margin
there (6.0 min vs 28.4) is not in question.

Domination (faster **and** better) on PSNR / on PSNR+SSIM:
[Brush](https://github.com/ArthurBrussee/brush) **48/48 · 48/48**,
[spirula-studio](https://github.com/harry7557558/spirula-studio) **45/48 · 41/48**,
[msplat](https://github.com/rayanht/msplat) **66/96 · 50/96**.

## ⏱️ Watch it converge

Four trainers, same seed cloud, each given the **same ~390 s**, running however many iterations fit.

![Side-by-side convergence against wall-clock](https://raw.githubusercontent.com/nandometzger/metal-gauss/main/assets/timelapse.gif)

| | final PSNR | first 20 dB | first 24 dB | first 27 dB |
|---|---:|---:|---:|---:|
| **metal-gauss** (15 k it) | **33.26** | 19.4 s | **57.5 s** | **147.1 s** |
| Brush (7 k it) | 24.82 | 105.6 s | 270.0 s | — |
| spirula-studio (5.5 k it) | 24.31 | 147.8 s | 371.6 s | — |
| msplat (19 k it) | 23.33 | **7.9 s** | 58.5 s | — |

lego, panel at 400 px, metrics over 20 held-out views at 800 px. Build the interactive version with
`python bench/compare/build_timelapse_page.py`.

## 📦 Install

```bash
pip install "git+https://github.com/nandometzger/metal-gauss"
```

macOS on Apple Silicon, Python ≥3.10, PyTorch ≥2.5. Metal kernels compile at **runtime** — no Xcode,
no `.metallib` step.

## 🚀 Train

```bash
metal-gauss-train --colmap scene/sparse/0 --images scene/images --steps 7000 --export scene.ply
metal-gauss-train --blender data/nerf_synthetic/lego --steps 30000
```

Every 8th view is held out. The exported `.ply` is standard INRIA-convention 3DGS. Capacity follows
the step count; `--budget` overrides it. Blender scenes are **not vendored** — unpack
[nerf_synthetic](https://github.com/bmild/nerf) into `data/nerf_synthetic/`.

## 🎥 Render

```bash
metal-gauss-render prediction.ply --out frame0.png --still --like-photo portrait.jpg
metal-gauss-render prediction.ply --out wiggle.mp4 --like-photo portrait.jpg
metal-gauss-render scene.ply --out orbit.mp4 --frame bbox --path orbit --sweep-deg 20
metal-gauss-render lego.ply --out lego.mp4 --frame bbox --up +z --path orbit --sweep-deg 60
```

Renders an existing `.ply` along a camera path the tool generates itself, so a file is no longer tied
to the dataset it came from. Training already wrote `.ply` files, but every render path in the repo
borrowed its cameras from a dataset, which left no way to look at a checkpoint, a scene trained
elsewhere, a download, or anything a feedforward model predicted. Point this at a file and get a png
or an mp4: `--path` picks still, wiggle or orbit, `--frame` picks what that path is built around,
and `--resolution`, `--fov`, `--background` and `--convention` do what they say.

`--frame input` anchors on the predicting camera, because a monocular predictor works in the input
photograph's frame and the identity world-to-camera matrix reproduces that shot. `--frame bbox`
places the camera around the cloud instead, for a trained scene that has no input camera, level
with it and orbiting its vertical axis. `--up` says which axis that is: the default `-y` is the
OpenCV world, and a scene trained with `--blender` keeps Blender's Z-up world, so it needs
`--up +z` or it is seen from underneath (`--up=-z` for a negative axis). The
default `auto` picks by the fraction of splats sitting in front of the origin, taking `input` above
**99%**.

`--like-photo` matters more than it looks, and it is the flag that makes a monocular prediction come
out right. A `.ply` carries no camera, but the prediction was made under one, taken from the
photograph's EXIF. Render it back at some other FOV and the geometry is right while the crop is not,
so frame 0 stops reproducing the photograph, which is the whole point of anchoring to it. Worked
example: Apple's [SHARP](https://github.com/apple/ml-sharp) reads `FocalLengthIn35mmFilm` and
converts it with `f_px = f_35mm · diag(W, H) / diag(36, 24)`; this flag reproduces that conversion
rather than approximating it, because the goal is to agree with the predictor and not to be
independently correct about the lens. Without the flag the FOV is fitted to the cloud and the tool
says so. The gap is not subtle: a 135 mm portrait is **12.9°**, SHARP's no-EXIF fallback of 30 mm
is **54°**. SHARP is also a fair test of the whole entry point, since its own video renderer
**requires CUDA** — predict on MPS, render here.

`--aperture` renders through a thin lens instead of a pinhole: the frame becomes the mean of many
views spread over the lens area, every one aimed at the focal plane, so that plane stays sharp and
everything else disperses. Defocus therefore needs no new rasteriser, only more cameras. Radius 0 is
the default and is bit-identical to the pinhole path. Below about 32 samples the sampling disc shows
as a lattice in the bokeh.

```bash
metal-gauss-render prediction.ply --out bokeh.mp4 --like-photo portrait.jpg \
    --aperture 0.03 --aperture-samples 96
```

Forward-only, so it is much faster than a training step: **69 fps** at 600k splats and 768², **91**
at 512², **481** at 100k splats and 384², on an M5 with the GPU to itself (`bench/render_fps.py`,
three round-robin repeats agreeing to within 10%). A defocused frame costs that times the sample
count.

Defaults are 60 frames at 30 fps, 512 px square, ±5° wiggle; `ffmpeg` on PATH writes the mp4.
`--still` dumps frame 0 as a `.png`, the cheap way to check a file's convention before rendering 60
frames of it; `--convention opengl` if it comes out flipped. A monocular prediction only has
evidence for what the photograph saw, so past roughly **8°** the sweep starts showing invented
surface. That is why the default sweep is small.

## 👀 View

```bash
pip install "metal-gauss[viewer] @ git+https://github.com/nandometzger/metal-gauss"
metal-gauss-view scene.ply --up +z
```

Open `http://127.0.0.1:8080`, drag to orbit and scroll to dolly. Every view is rendered on the Mac's
GPU and sent to the browser tab as an image, so the browser does no splatting of its own and the
viewer shows exactly what `metal-gauss-render` would. While the camera moves, a frame too slow for
30 fps is rendered at a lower resolution. Once the camera stops, it gets one full-resolution frame,
and after that nothing, so an idle viewer leaves the GPU alone.

The Lens panel is the same thin lens as `--aperture`. It refines while the camera is still, showing
the running mean at 8, 16, 32 and 64 samples and then at all of them. `--up`, `--convention`, `--fov`
and `--background` mean what they do for `metal-gauss-render`.

The Crop panel cuts away what you do not want: drag the box, rotate it, or set its size, and
everything outside it disappears from the view, from an exported video, and from **Export cropped
.ply**. A short run leaves a haze of near-transparent floaters around the model, and this is what
removes them from the file rather than merely framing around them. Training is never affected.

The Path panel flies a camera through the scene and writes an mp4. **Orbit** and **Wiggle** lay
down keyframes around the current view; **Add keyframe** captures wherever you are, and the path
runs smoothly through them and loops. Frames, fps and resolution are yours to set, and the current
aperture can come along for a defocused video. Rendering happens here rather than in the browser,
so the video comes out at full resolution with the depth of field intact.

The server listens on localhost only. To view from another machine, tunnel it
(`ssh -L 8080:127.0.0.1:8080 your-mac`) or pass `--host 0.0.0.0` to serve it to the network.
[viser](https://github.com/nerfstudio-project/viser) is an optional extra, so the base install
does not pull it in. Keep the tab visible: a hidden one stops the browser's render loop, and the
page goes blank until you come back to it.

### Watch it train

```bash
metal-gauss-train --blender data/nerf_synthetic/lego --steps 7000 --viewer
```

The same page shows the model while it trains. It opens on the first training camera, with the up
direction taken from the cameras themselves. It is rendered between optimisation steps, on the
training thread, and may take at most `--viewer-budget` of wall-clock (10% by default, adjustable
from the page). The budget also pays for the GPU syncs the preview forces.

On lego, 2000 steps, in interleaved runs:
- **No browser connected:** no measurable cost (−1.2%, against 3.4% run-to-run noise).
- **One browser connected:** +5.4% ms/step on means and +8.6% on medians, against about 10% noise.

- **Training panel:** step, loss, held-out PSNR, an ETA, and the preview's share of wall-clock. The
  ETA prices the curriculum that is left — the resolution schedule, the capacity ramp and the evals
  still to come — so it does not read low early. It is in the terminal log too, with or without the
  viewer. **Pause**
  gives the preview the whole GPU, lens included; paused time is left out of every reported time.
- **Cameras:** the training cameras (blue) and held-out cameras (orange). Held-out cameras appear
  once that split has been loaded for evaluation. Click one to jump to it.
- **Snapped view:** shows that camera's PSNR, and **Show photo** swaps the render for its
  photograph, pixel-aligned.
- **When training ends,** the page keeps serving the final model until Ctrl-C.

Exporting a video while training shares the same budget, so it renders slowly and the model keeps
improving between frames — a progress flythrough rather than a turntable. Pause first for a
consistent one, which also runs at full speed.

`--viewer` is off by default. The benchmark harness refuses reports from runs that used it, because
the preview shares the GPU.

## ⚠️ Caveats

- msplat is **1.3–1.8× faster per step**. Our fixed startup at that end was ~8.4 s and is now
  **~3 s**: 6.1 s of it was decoding PNGs, and 4.1 s of that decoded the 200 held-out views a run
  without evaluation never reads. The metal-gauss rows above were re-measured after that change;
  the competitor rows were not, and predate it.
- That re-measurement moves the **1 min** budget by +0.5 dB and nothing else — and **do not lean on
  that +0.5 dB**. It comes from three scenes whose 2 000-iteration rung cleared 60 s by 0.5, 0.9 and
  2.2 s; re-measured later the same day against a machine also driving a desktop, none of the eight
  cleared it. The rung boundary, not the loader, decides that row. What is solid is the saving
  itself: ~4 s, visible at the 500 rung (median −4.4 s over 8 scenes, every scene negative) and
  invisible past ~2 000, where run-to-run spread is ±40 s.
- Wall-clock rows need a machine that is not also drawing a screen. The same 2 000-iteration rung
  measured 57–68 s on an idle machine and 74–96 s while a browser was open. `require_gpu_exclusive()`
  catches a second trainer, not a compositor.
- 7 k numbers are **not comparable to published 30 k numbers** — 5.1 dB apart.
- `--antialias` is off by default; worth **+6.68 dB at 200 px** render resolution.
- Run-to-run noise floors: ours **0.19 dB**, Brush **0.74**, spirula **0.15–1.27**, msplat **3.35**
  (no seed flag).
- `--budget` at 1 M splats: **+2.75 dB** lego, **+1.40** mic, **−0.20** ship, all at ~5× the time.

## 📚 More

| | |
|---|---|
| [docs/BENCHMARKS.md](https://github.com/nandometzger/metal-gauss/blob/main/docs/BENCHMARKS.md) | full results, protocol, calibration, noise floors |
| [docs/ARCHITECTURE.md](https://github.com/nandometzger/metal-gauss/blob/main/docs/ARCHITECTURE.md) | how it works, the kernels, correctness oracle |
| [bench/results/NEGATIVE_RESULTS.md](https://github.com/nandometzger/metal-gauss/blob/main/bench/results/NEGATIVE_RESULTS.md) | **every rejected lever and measurement lesson, with numbers** |
| [bench/compare/STATUS.md](https://github.com/nandometzger/metal-gauss/blob/main/bench/compare/STATUS.md) | every Apple-native implementation surveyed, and the traps |

`NEGATIVE_RESULTS.md` is the most useful file here: several published speedups measure near zero on
this hardware, and several of this repo's own conclusions were wrong until re-measured.

## 📄 License

MIT. Credits: [3DGS-MCMC](https://arxiv.org/abs/2404.09591),
[Taming 3DGS](https://arxiv.org/abs/2406.15643), [Speedy-Splat](https://arxiv.org/abs/2412.00578),
[LiteGS](https://arxiv.org/abs/2503.01199), [Mip-Splatting](https://arxiv.org/abs/2311.16493),
[gsplat](https://github.com/nerfstudio-project/gsplat),
[LichtFeld Studio](https://github.com/MrNeRF/LichtFeld-Studio),
[Brush](https://github.com/ArthurBrussee/brush), and the original
[INRIA rasterizer](https://github.com/graphdeco-inria/gaussian-splatting).
