Metadata-Version: 2.5
Name: kintsugi-st
Version: 0.5.0
Summary: Tissue-wide RNA density and composition estimation from thinned molecule splits
Project-URL: Homepage, https://github.com/cafferychen777/kintsugi
Project-URL: Repository, https://github.com/cafferychen777/kintsugi
Project-URL: Issues, https://github.com/cafferychen777/kintsugi/issues
Project-URL: Documentation, https://github.com/cafferychen777/kintsugi#readme
Author: Kintsugi authors
Maintainer: Kintsugi authors
License: MIT License
        
        Copyright (c) 2026 Kintsugi authors
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: Poisson thinning,density estimation,spatial transcriptomics,visium hd
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.10
Requires-Dist: anndata>=0.10
Requires-Dist: h5py>=3.10
Requires-Dist: numpy>=1.24
Requires-Dist: pandas>=2.0
Requires-Dist: pyarrow>=14.0
Requires-Dist: scipy>=1.11
Provides-Extra: dev
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Description-Content-Type: text/markdown

# Kintsugi

Kintsugi turns binned spatial transcriptomics counts (for example 8 µm
Visium HD bins) into two fields defined on every tissue bin, including bins
that captured no molecule:

| Field | Meaning |
| --- | --- |
| `density` | Captured molecules per bin at unit thinning exposure; divide by bin area for an areal rate |
| composition | Per-bin gene probabilities that sum to one; density times composition is the expected count of each gene |

The fields are predictions, not observed counts. Each gene borrows
information from a validation-selected mixture of pooled estimates: square
grids, connected regions found by a joint Poisson merge of all genes and their
boundary-refined versions, Gaussian kernels, and a low-rank factor family in
which genes that follow shared programmes borrow strength across the whole
section. Gene-wise weights let every gene choose its own scale, so a sharp
marker and a diffuse housekeeping gene are pooled differently.

Nothing is tuned by hand. The molecules are thinned once into three
independent splits: A fits every candidate, B selects among them, and C is an
untouched held-out split for scoring. Geometry is built from A only.

## Install

```bash
python -m pip install kintsugi-st
```

The distribution is `kintsugi-st`; the import is `kintsugi`. Do not install
the unrelated PyPI package named `kintsugi`. From a checkout,
`python -m pip install -e ".[dev]"` adds pytest and ruff. Python 3.10 or later;
NumPy, SciPy, h5py, pandas, PyArrow and AnnData are pulled in; no GPU.

## The published estimator in one call

`estimate_fields` is the estimator used in the manuscript experiments
(same-section Visium HD versus Xenium, Xenium Protein RCC, Vannan PF cohort).
It builds the A-only candidate bank, fits and stacks the components with
gene-wise weights, adds the factor family with prior-shrunk corrections, selects
everything on B and refits the selected components on A+B. C is never read.

From a Space Ranger output directory

```text
sample/
├── filtered_feature_bc_matrix.h5
└── spatial/
    └── tissue_positions.parquet
```

the complete workflow on one section is:

```python
import numpy as np
from kintsugi import (
    load_visium_hd_fields, section_fit_inputs, estimate_fields,
    save_fields, load_fields, score_fields,
)

# 1. Load. Gene Expression features only, duplicate symbols summed, every
#    in-tissue bin retained. data.counts is sparse (n_bins, n_genes) with rows
#    in np.where(data.mask) order; data.mask is the (H, W) bin grid.
data = load_visium_hd_fields(
    "sample/filtered_feature_bc_matrix.h5",
    "sample/spatial/tissue_positions.parquet",
    bin_um=8, dataset="sample",
)

# 2. Thin every molecule once, section-wide, into A (50%), B (25%), C (25%),
#    and fit the section-wide factor loadings on A. rows=None keeps all bins;
#    pass the row indices of a window to fit only part of the section.
A, B, C, loadings = section_fit_inputs(data.counts, data.mask, rows=None, seed=42)

# 3. Fit on A, select on B, refit the selection on A+B.
fit = estimate_fields(A, B, data.mask, loadings=loadings)

# 4. Save and reload (a new directory); names guard against gene-order mistakes.
save_fields(fit, "sample_fields", gene_names=data.gene_names)
fit = load_fields("sample_fields", expected_mask=data.mask,
                  expected_gene_names=data.gene_names)

# 5. Score on the untouched C. Do this once, after every B choice is final.
metrics = score_fields(fit, C, exposure=0.25, mask=data.mask,
                       gene_names=data.gene_names)
print(metrics["joint_point_logscore"], metrics["heldout_UMI"])
```

`fit.density` is one value per tissue bin in mask order. The composition is
streamed rather than stored: `fit.composition_chunks(chunk_size=64)` yields
consecutive gene blocks of shape `(n_bins, chunk)`, and `intensity_chunks`
yields density times composition. The probability that a molecule at a bin
belongs to a gene programme is the composition mass on those genes:

```python
genes = list(data.gene_names)
columns = np.array([genes.index(g) for g in ("COL1A1", "COL1A2", "COL3A1")])

from kintsugi import programme_probability

image = np.full(data.mask.shape, np.nan)
image[data.mask] = programme_probability(fit, columns)
```

`fit.density_specification["name"]` names the selected density pooling,
`fit.weights` are the mixture weights (predictive weights, not cell-type
proportions) and `fit.optimizer` records every B choice.

Memory grows with the number of bins and the candidate bank. Windows of about
10,000 bins fit comfortably; whole sections go through the tiled driver.

## Whole sections

`fit_fields_tiled` runs the same estimator on overlapping square tiles and
stitches the result, each bin estimated once by the tile whose centre is
nearest. The factor loadings are fitted once on the section-wide A, so all
tiles share one programme basis.

```python
from kintsugi import split_counts, fit_fields_tiled, programme_probability

A, B, C = split_counts(data.counts, seed=42)
tiled = fit_fields_tiled(A, B, data.mask, tile=128, step=96)

tiled.density                          # (n_bins,) in mask order
for start, block in tiled.composition_chunks(64):
    ...                                # stitched gene blocks
prob = programme_probability(tiled, columns)
np.save("sample_collagen_probability.npy", prob)
```

The tiled result has no `save_fields` form; save the density and the
programme probabilities you need with NumPy. Tiles fit independently, so a
machine with a few tens of gigabytes handles a full Visium HD section given
time.

## Evaluation

`kintsugi.evaluate` holds the positional endpoints of the experiments as pure
NumPy functions on per-bin vectors: `signed_distance` to a reference
compartment boundary with censoring near the field of view, `edge_zone` and
`auc` for boundary discrimination, `contrast_ratio` and `compartment_means`
between compartments, `area_matched_mask` with `boundary_metrics` for
area-matched boundary agreement, `band_profile` and `band_nmse` for distance
profiles, `programme_fraction`, `programme_binomial_score` and
`per_molecule_logscore` for held-out programme scores, and
`positional_endpoints` combining them for one programme against a reference
such as same-section Xenium.

## Options

`estimate_fields` forwards any keyword to `fit_fields`, and each option is
still selected on B. The defaults are the published preset; these are the
knobs that change the estimator:

- `gene_tau=True` selects a per-gene shrinkage inside every plain pooled spatial component (the factor-informed corrections have their own switch, below). It improves the all-gene held-out likelihood but can reduce programme contrast; use it when per-gene prediction matters more than sharp programme maps.
- `lowrank_prior_gene_tau=True` lets each factor-informed correction choose its shrinkage per gene rather than one for all genes, so how far a gene's own neighbourhood overrides the factor prediction depends on the gene. It raises the validation likelihood of every correction substantially, but the held-out gain does not follow: it helps on shallow panels with an image reference and costs placement on the 5,000-gene panel. Off by default; see research/genewise_tau_20260909.
- `lowrank_prior="multiplicative"` replaces the additive (Dirichlet-multinomial) factor-informed correction with the renormalised Gamma-Poisson form. Keep the default unless you are comparing the two forms.
- `lowrank_prune=True` keeps only the factor components and their prior-shrunk corrections, the compact form of the estimator. It is not the default because with small panels (about twenty genes) the factor prior is poor and the plain spatial components are the safety net.
- `refit="ABC"` with `test_counts=C` rebuilds the selection on all three splits for a final field once scoring is finished; such a fit carries `optimizer["refit"]["test_used"]` and `score_fields` refuses it.

For an explicit bank, `build_candidate_bank(A, mask)` returns the published
candidates and `fit_fields(A, B, mask, candidates, ...)` is the long form of
`estimate_fields`. `split_counts_rows` returns the window rows of a
section-wide split without the loadings.

## Legacy API (0.1.x)

The 0.1.x total-count pipeline (variogram guide, seeded watershed
tessellation, tessellation report, the earlier candidate menu) was superseded
by the joint-merge geometry and is no longer exported from `kintsugi`. It
remains importable from `kintsugi.legacy`:

```python
import kintsugi
import kintsugi.legacy

grid = kintsugi.load_visium_hd_from_dir("sample/")
result = kintsugi.legacy.tessellate(grid)
print(kintsugi.legacy.tessellation_report(result, grid))
adata = kintsugi.to_anndata(result, grid=grid, use_raw_counts=True)
```

`kintsugi.legacy.adaptive_tessellation`, `directional_semivariance`,
`boundary_tensor` and `iter_pooling_candidates` are there as well.
`aggregate_counts`, `build_spatial_graph` and `to_anndata` stay in `kintsugi`
for regional count summaries of any connected labelling. Runnable examples of
both the estimator and the legacy pipeline are in `examples/`.

## Development checks

```bash
python -m ruff check kintsugi tests
python -m pytest -q
```

## Citation

If you use Kintsugi in your research, please cite the associated manuscript when
it becomes available; `CITATION.cff` carries the software metadata.

## License

Kintsugi is released under the [MIT License](LICENSE).
