Metadata-Version: 2.3
Name: histotuner
Version: 0.3.3
Summary: Add your description here
Author: Ajit Johnson Nirmal
Author-email: Ajit Johnson Nirmal <ajitjohnson.n@gmail.com>
Requires-Dist: anndata>=0.12.2
Requires-Dist: cellpose>=4.0.6
Requires-Dist: dask>=2024.11.2
Requires-Dist: geopandas>=1.1.1
Requires-Dist: leidenalg>=0.12.0
Requires-Dist: magicgui>=0.10.1
Requires-Dist: matplotlib>=3.10.6
Requires-Dist: napari>=0.7.0
Requires-Dist: numpy>=2.3.3
Requires-Dist: opencv-python>=4.11.0.86
Requires-Dist: openslide-bin>=4.0.0.8
Requires-Dist: openslide-python>=1.4.2
Requires-Dist: pandas>=2.3.3
Requires-Dist: pillow>=11.3.0
Requires-Dist: pip>=25.2
Requires-Dist: psutil>=7.1.0
Requires-Dist: pyqt6>=6.11.0
Requires-Dist: python-igraph>=1.0.0
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: scikit-learn>=1.7.2
Requires-Dist: shapely>=2.1.2
Requires-Dist: spatialdata>=0.5.0
Requires-Dist: tifffile>=2025.9.30
Requires-Dist: timm>=1.0.20
Requires-Dist: tqdm>=4.67.1
Requires-Dist: transformers>=4.57.1
Requires-Dist: umap-learn>=0.5.7
Requires-Dist: wandb>=0.22.2
Requires-Dist: zarr>=3
Requires-Dist: cupy-cuda12x ; extra == 'linux-gpu'
Requires-Dist: cudf-cu12 ; extra == 'linux-gpu'
Requires-Dist: cugraph-cu12 ; extra == 'linux-gpu'
Requires-Dist: cuml-cu12 ; extra == 'linux-gpu'
Requires-Python: >=3.12
Provides-Extra: linux-gpu
Description-Content-Type: text/markdown

## histotuner

### GPU UMAP and Clustering on Linux

`histotuner` installs CPU UMAP support through `umap-learn`. GPU UMAP is
optional because it depends on the local CUDA driver/toolkit stack and should be
installed separately from the package dependencies in `pyproject.toml`.

GPU UMAP uses RAPIDS cuML when both `cuml` and `cupy` are available in the
active Python environment and a CUDA-capable NVIDIA GPU is visible.

The same optional GPU stack is also used by native clustering:

- `ht.umap(...)`
- `ht.leiden(...)`
- `ht.dbscan(...)`
- `histotuner-leiden`
- `histotuner-dbscan`

Check the active environment from Python:

```python
import histotuner as ht

ht.umap_backend_status()
```

Expected GPU-ready output has `gpu_available: True`, with both `gpu_cuml` and
`gpu_cupy` set to `True`.

Recommended install path on Linux is to create a RAPIDS-compatible environment
with the official RAPIDS install selector:

https://docs.rapids.ai/install/

For conda/mamba environments, install at least `cuml` and the matching CUDA
runtime package for your machine. A typical CUDA 12-style command looks like:

```bash
mamba create -n histotuner-rapids \
  -c rapidsai -c conda-forge -c nvidia \
  python=3.12 cuml cuda-version=12.0

mamba activate histotuner-rapids
pip install -e .
```

For pip-based RAPIDS installs, choose wheels matching the installed CUDA major
version. `histotuner` now provides an optional extra for a CUDA 12 Linux GPU
stack:

```bash
pip install -e ".[linux-gpu]" --extra-index-url=https://pypi.nvidia.com
```

That extra currently expands to:

```bash
pip install cupy-cuda12x cudf-cu12 cugraph-cu12 cuml-cu12 \
  --extra-index-url=https://pypi.nvidia.com
```

If your Linux environment is not CUDA 12 based, use the RAPIDS selector to
generate the correct command for that exact CUDA/RAPIDS combination instead of
the `linux-gpu` extra.

Then run UMAP with:

```python
ht.umap(
    sdata=zarr_path,
    tableKeys=["mstar_tokens", "virchow2_tokens"],
    sample_n=25000,
    prefer_gpu="auto",  # uses GPU if RAPIDS is available, otherwise CPU
)
```

To require GPU and fail loudly if RAPIDS is not available:

```python
ht.umap(
    sdata=zarr_path,
    tableKeys=["mstar_tokens", "virchow2_tokens"],
    sample_n=25000,
    prefer_gpu="gpu",
)
```

Native clustering uses the same `prefer_gpu` switch:

```python
ht.leiden(
    sdata=zarr_path,
    tableKeys="tokens",
    obsm_key="X_umap",
    prefer_gpu="auto",
    target_col="leiden",
)
```

```python
ht.dbscan(
    sdata=zarr_path,
    tableKeys="tokens",
    obsm_key="X_umap",
    prefer_gpu="gpu",
    target_col="dbscan",
)
```

CLI examples:

```bash
histotuner-leiden /path/to/sample.zarr \
  --tables tokens \
  --obsm-key X_umap \
  --prefer-gpu auto \
  --target-col leiden
```

```bash
histotuner-dbscan /path/to/sample.zarr \
  --tables tokens \
  --obsm-key X_umap \
  --prefer-gpu gpu \
  --target-col dbscan
```

Notes:

- RAPIDS requires Linux or WSL2; native Windows Python environments generally
  cannot install/use cuML, cuGraph, or cuDF directly.
- GPU UMAP uses `cuml` plus `cupy`.
- GPU Leiden uses `cudf`, `cugraph`, `cuml`, and `cupy`.
- GPU DBSCAN uses `cuml` plus `cupy`.
- CUDA package suffixes must match the CUDA toolkit/driver stack in the
  environment. If installation fails, generate a fresh command from the RAPIDS
  selector for the specific Linux, Python, CUDA, and RAPIDS versions.

### Supported token-extraction backends

`histotuner` can append multiple model-specific token tables into the same
SpatialData Zarr while keeping shared geometry layers model-agnostic.

Currently supported token extractors:

- `hf-hub:bioptimus/H-optimus-1`
- `hf-hub:MahmoodLab/UNI2-h`
- `hf-hub:paige-ai/Virchow2`
- `hf-hub:Wangyh/mSTAR`
- `hf-hub:prov-gigapath/prov-gigapath`
- `owkin/phikon-v2`
- `MahmoodLab/conchv1_5`
- `WenchuanZhang/Patho-CLIP-L`
- `majiabo/GPFM`
- `kaiko-ai/vitl14`
- `xiangjx/musk`

### Token-grid semantics

All currently supported models export a unified `14x14` token grid so token
tables can be compared directly across models.

- `phikon-v2` exports a native `14x14` patch-token grid.
- `hf-hub:bioptimus/H-optimus-1`, `hf-hub:Wangyh/mSTAR`, and
  `hf-hub:prov-gigapath/prov-gigapath` export native `14x14` grids.
- `hf-hub:MahmoodLab/UNI2-h` and `hf-hub:paige-ai/Virchow2` have native
  `16x16` patch-token grids after special tokens are stripped, and `histotuner`
  adaptively average-pools them to `14x14`.
- `conchv1_5` is special:
  - the native vision encoder runs at `448x448` with `patch16`
  - that produces a native `28x28` patch-token grid
  - `histotuner` average-pools each non-overlapping `2x2` token neighborhood
    to export a compatibility `14x14` token grid
- `Patho-CLIP-L` is also special:
  - the native CLIP-L/14 vision encoder produces a `24x24` patch-token grid at
    `336x336` input resolution
  - `histotuner` adaptively average-pools that native `24x24` grid to export a
    compatibility `14x14` token grid
- `GPFM` is also special:
  - the native DINOv2 ViT-L/14 encoder produces a `16x16` patch-token grid at
    `224x224` input resolution
  - `histotuner` adaptively average-pools that native `16x16` grid to export a
    compatibility `14x14` token grid
- `kaiko-ai/vitl14` is also special:
  - the native Kaiko ViT-L/14 encoder produces a `16x16` patch-token grid at
    `224x224` input resolution
  - `histotuner` uses the Kaiko preprocessing defaults (`mean=std=0.5`) and
    adaptively average-pools that native `16x16` grid to export a compatibility
    `14x14` token grid
- `xiangjx/musk` is also special:
  - the native MUSK patch16 vision encoder produces a `24x24` patch-token grid
    at `384x384` input resolution
  - `histotuner` uses the MUSK preprocessing defaults (`mean=std=0.5`) and
    adaptively average-pools that native `24x24` grid to export a compatibility
    `14x14` token grid
  - MUSK is gated on Hugging Face and requires the optional official `musk`
    package

That pooling choice is deliberate so downstream single-cell workflows can
consume every supported model through the same `14x14` token layout. For the
pooled models, this is a compatibility semantic rather than the model's native
tokenization:

- `UNI2-h` and `Virchow2`: pooled from native `16x16`
- `conchv1_5`: pooled from native `28x28`
- `Patho-CLIP-L`: pooled from native `24x24`
- `GPFM`: pooled from native `16x16`
- `kaiko-ai/vitl14`: pooled from native `16x16`
- `xiangjx/musk`: pooled from native `24x24`

### Not yet supported for token extraction

- none from the current requested set

### O2 batch job generation

To generate one `embedder.yaml` and one `embed_cluster.sh` per sample folder on
O2:

```bash
python generate_o2_jobs.py \
  --root-dir /n/scratch/users/a/ajn16/histotuner/full \
  --template-yaml embedder.yaml \
  --template-shell embed_cluster.sh \
  --output-dir /n/scratch/users/a/ajn16/histotuner/generated_jobs

  python generate_o2_jobs.py \
  --root-dir /n/scratch/users/a/ajn16/histotuner/heonly \
  --template-yaml embedder_HEonly.yaml \
  --template-shell embed_cluster_HEonly.sh \
  --output-dir /n/scratch/users/a/ajn16/histotuner/generated_jobs


```

To preview the `sbatch` submissions for the generated job scripts:

```bash
python submit_generated_jobs.py \
  --generated-dir /n/scratch/users/a/ajn16/histotuner/generated_jobs \
  --dry-run
```

### Melanocyte UMAP/DBSCAN token pipeline

The scripts in `o2/melanocyte_dbscan/` find SpatialData `.zarr` stores under a
folder, map tokens to cells, compute UMAP for native token tables, map broad
phenotype labels onto token tables, and run DBSCAN on `X_umap` for tokens where
`phenotype_broad == "Melanocytes"`. They then generate thumbnail PDFs for
`dbscan_melanocytes_umap` using an HE image auto-detected beside each `.zarr`,
and save two UMAP plots for each token table/model:

- `dbscan_melanocytes_umap`, excluding `-1` and `nan`
- `phenotype_broad`, excluding `-1` and `0`

Run a local dry-run first:

```powershell
python .\o2\melanocyte_dbscan\run_token_umap_melanocyte_dbscan.py `
  "C:\Users\aj\Downloads\test" `
  --recursive `
  --dry-run
```

Run the local pipeline and write a summary:

```powershell
python -u .\o2\melanocyte_dbscan\run_token_umap_melanocyte_dbscan.py `
  "C:\Users\aj\Downloads\test" `
  --recursive `
  --continue-on-error `
  --summary-json "C:\Users\aj\Downloads\test\melanocyte_dbscan_summary.json"
```

The script uses native histotuner token-table selection. By default,
`tokenCellMapper` and `phenotypeCellMapper` auto-detect token tables, while UMAP
and DBSCAN use the native `tokens` selector. Melanocyte DBSCAN uses
`--dbscan-min-samples 100` by default. Pass `--no-thumbnail-pdfs` to skip PDF
generation, or `--thumbnail-image-path /path/to/image.ome.tiff` to provide an
explicit image for a single-sample run.
Pass `--no-umap-plots` to skip the saved UMAP plots.
The pipeline runs `ht.repairSpatialDataTableRegistry(...)` at the start, after
UMAP writes, and after DBSCAN writes so on-disk tables are re-registered before
downstream plotting/PDF steps.

### O2 parallel melanocyte DBSCAN jobs

On O2, generate one Slurm script per sample/zarr without submitting:

```bash
python o2/melanocyte_dbscan/submit_melanocyte_dbscan_jobs.py \
  --root-dir /n/scratch/users/a/ajn16/he_embed \
  --output-dir /n/scratch/users/a/ajn16/melanocyte_dbscan_jobs
```

Submit a single test job:

```bash
python o2/melanocyte_dbscan/submit_melanocyte_dbscan_jobs.py \
  --root-dir /n/scratch/users/a/ajn16/he_embed \
  --output-dir /n/scratch/users/a/ajn16/melanocyte_dbscan_jobs_test \
  --limit 1 \
  --submit
```

Submit all jobs:

```bash
python o2/melanocyte_dbscan/submit_melanocyte_dbscan_jobs.py \
  --root-dir /n/scratch/users/a/ajn16/he_embed \
  --output-dir /n/scratch/users/a/ajn16/melanocyte_dbscan_jobs \
  --submit
```

Replace `/n/scratch/users/a/ajn16/he_embed` with the O2 path containing the
sample folders or `.zarr` stores. Each submitted job runs the melanocyte DBSCAN
pipeline on one sample folder, so samples run in parallel through Slurm.
Thumbnail PDFs are generated by default in each sample job; pass
`--no-thumbnail-pdfs` to
`o2/melanocyte_dbscan/submit_melanocyte_dbscan_jobs.py` to disable them. UMAP
plots are also generated by default; pass `--no-umap-plots` to disable them.
The submission manifest is written to
`/n/scratch/users/a/ajn16/melanocyte_dbscan_jobs/submission_manifest.csv`.
