Metadata-Version: 2.3
Name: poincar3
Version: 1.0.0
Summary: Emergent Multi-View Geometry Through Self-Distillation.
Keywords: computer-vision,self-supervised-learning,multi-view,3d,vision-transformer
Author: David Nordström
Author-email: David Nordström <davnords@chalmers.se>
License: MIT License
         
         Copyright (c) 2026 David Nordström
         
         Permission is hereby granted, free of charge, to any person obtaining a copy
         of this software and associated documentation files (the "Software"), to deal
         in the Software without restriction, including without limitation the rights
         to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
         copies of the Software, and to permit persons to whom the Software is
         furnished to do so, subject to the following conditions:
         
         The above copyright notice and this permission notice shall be included in all
         copies or substantial portions of the Software.
         
         THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
         IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
         FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
         AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
         LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
         OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
         SOFTWARE.
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Dist: numpy>=1.26
Requires-Dist: torch
Requires-Dist: opencv-python>=4.12.0.88 ; extra == 'demo'
Requires-Dist: pillow>=10.0.0 ; extra == 'demo'
Requires-Dist: tyro ; extra == 'demo'
Requires-Dist: poincar3[train] ; extra == 'eval'
Requires-Dist: flow-vis>=0.1 ; extra == 'eval'
Requires-Dist: kornia>=0.8.2 ; extra == 'eval'
Requires-Dist: loguru>=0.7.0 ; extra == 'eval'
Requires-Dist: openexr>=3.4.13 ; extra == 'eval'
Requires-Dist: romatch>=0.1.2 ; extra == 'eval'
Requires-Dist: poincar3[demo] ; extra == 'train'
Requires-Dist: h5py>=3.12.1 ; extra == 'train'
Requires-Dist: matplotlib>=3.10.7 ; extra == 'train'
Requires-Dist: rich>=14.2.0 ; extra == 'train'
Requires-Dist: scipy>=1.15.3 ; extra == 'train'
Requires-Dist: torchvision>=0.23.0 ; extra == 'train'
Requires-Dist: tqdm>=4.67.1 ; extra == 'train'
Requires-Dist: wandb>=0.23.0 ; extra == 'train'
Requires-Python: >=3.10
Project-URL: Homepage, https://github.com/davnords/poincar3
Project-URL: Repository, https://github.com/davnords/poincar3
Provides-Extra: demo
Provides-Extra: eval
Provides-Extra: train
Description-Content-Type: text/markdown

<div align="center">
<h1>Emergent Multi-View Geometry Through Self-Distillation</h1>

<a href="https://arxiv.org/abs/XXXX.XXXXX"><img src="https://img.shields.io/badge/arXiv-XXXX.XXXXX-b31b1b" alt="arXiv"></a>
<a href="https://pypi.org/project/poincar3/"><img src="https://img.shields.io/pypi/v/poincar3" alt="PyPI"></a>
<a href="https://www.davnords.com/poincar3"><img src="https://img.shields.io/badge/Project_Page-green" alt="Project Page"></a>

[David Nordström<sup>1</sup>](https://scholar.google.com/citations?user=-vJPE04AAAAJ), [Thibaut Loiseau<sup>2</sup>](https://scholar.google.com/citations?user=qDSlhTUAAAAJ), [Vincent Lepetit<sup>2</sup>](https://scholar.google.com/citations?user=h0a5q3QAAAAJ&hl),<br>[Michael Felsberg<sup>3</sup>](https://scholar.google.com/citations?user=lkWfR08AAAAJ), [Guillaume Bourmaud<sup>4</sup>](https://scholar.google.com/citations?user=d4v2IYMAAAAJ), [Fredrik Kahl<sup>1</sup>](https://scholar.google.com/citations?user=P_w6UgMAAAAJ)

<sup>1</sup> **Chalmers University of Technology**, <sup>2</sup> **Ecole Nationale des Ponts et Chaussées, IP Paris**,<br><sup>3</sup> **Linköping University**, <sup>4</sup> **University of Bordeaux, CNRS**
</div>

<p align="center">
    <img src="assets/demo.png" alt="Emerging matching capabilities from self-supervision." width=64%>
    <img src="assets/poincare.png" alt="Henri Poincaré" width=17%>
    <br>
    <em>Left: Emerging matching capabilities from self-supervision. Simply using Poincar3's attention map reliable tracks can be created. Produced by <code>demo.py</code>. Right: Henri Poincaré argued that a being without motion cannot understand 3D space.</em>
</p>

## Overview

We release a model, Poincar3, that learns multi-view geometry from only training on image sequences. Poincar3 uses a multi-view transformer and a self-distillation objective giving strong zero-shot features and attention maps.

## Install

```bash
uv add poincar3          # or: pip install poincar3
```

The model itself needs only `torch`. Extras pull in what the scripts need:

```bash
uv add "poincar3[demo]"     # demo.py
uv add "poincar3[train]"    # training, mvcorr and SeeSE3 evals
uv add "poincar3[eval]"     # + matchbench and feed-forward reconstruction
```

To work from a clone instead (tested on Linux with Python 3.12):

```bash
uv sync --extra eval
```

## Usage

The checkpoint auto-downloads on first use:

```python
import torch
from poincar3 import Poincar3

model = Poincar3().eval().cuda()

# A multi-view batch: [batch, frames, 3, H, W], RGB in [0, 1], H and W multiples of 16.
images = torch.rand(1, 4, 3, 448, 448).cuda()

with torch.no_grad():
    patch_logits, patch_features, global_logits, camera_tokens = model(images)

# patch_features:  [1, 4, 784, 1024]  dense per-frame tokens, cross-view attended
# camera_tokens:   [1, 4, 1024]       one per-frame scene/pose token
```

### Demo

We illustrate how to create attention tracks by simply running:
```bash
uv run python demo.py
```
We also provide code for plotting the raw feature correlations in `cross_correlations.py`.

## Pretrained weights

Auto-downloaded on first use. To fetch it directly:

- **Poincar3 ViT-L/16** — [Download](https://github.com/davnords/storage/releases/download/mumv2/poincar3.pth)

## Evaluation

All evaluations, except feed-forward reconstruction, can be accessed through the `experiments/eval.py` endpoint. You can get ScanNet and NAVI following [these](https://github.com/mbanani/probe3d/blob/main/data_processing/README.md) instructions. To get the possible configurations, simply run it with the flag --help. For example, you run multi-view correspondence estimation on scannet by:
```bash
uv run python experiments/eval.py --evaluation mvcorr --mvcorr.dataset scannet
```

### Feed-forward reconstruction

Camera-pose and per-pixel-depth heads on top of the backbone, under two protocols:

```bash
# full finetune
torchrun --nproc-per-node 4 experiments/ffrecon/train.py --name finetune-poincar3

# frozen backbone + a small adapter
torchrun --nproc-per-node 4 experiments/ffrecon/train_adapter.py --backbone poincar3

# relative pose evaluation
uv run python experiments/ffrecon/eval.py --checkpoint <path> --evaluation relpose --relpose.dataset megadepth
# point-cloud estimation
uv run python experiments/ffrecon/eval.py --checkpoint <path> --evaluation pointcloud --pointcloud.dataset eth3d
```

## Training

```bash
torchrun --nproc-per-node 4 experiments/train.py --name my-run
```

### Training data

While we trained on large collection of 3D datasets, we illustrate our training protocol on [ScanNet++](https://scannetpp.mlsg.cit.tum.de/scannetpp/) and [RealEstate10K](https://google.github.io/realestate10k/). You can download them using their official download links.  

## License

MIT, except where a file notes otherwise. `src/poincar3/layers/` and parts of `heads/` derive
from [DINOv3](https://github.com/facebookresearch/dinov3) and
[VGGT](https://github.com/facebookresearch/vggt) and carry their original licenses;
`benchmarks/mv_consistency/` is adapted from [probe3d](https://github.com/mbanani/probe3d) (MIT).

## Acknowledgement

Built on [DINOv3](https://github.com/facebookresearch/dinov3),
[DINOv2](https://github.com/facebookresearch/dinov2),
[VGGT](https://github.com/facebookresearch/vggt),
[probe3d](https://github.com/mbanani/probe3d) and [RoMa](https://github.com/Parskatt/RoMa).

## BibTeX

```bibtex
TBD
```
