Metadata-Version: 2.4
Name: osteosarc
Version: 0.7.0
Summary: Reproducible access to the public osteosarc.com dataset
Author-email: Alex Rubinsteyn <alex.rubinsteyn@gmail.com>
License-Expression: Apache-2.0
Project-URL: Documentation, https://iskandr.github.io/osteosarc/
Project-URL: Source, https://github.com/iskandr/osteosarc
Project-URL: Issues, https://github.com/iskandr/osteosarc/issues
Project-URL: Dataset, https://osteosarc.com/data/
Keywords: osteosarcoma,genomics,neoantigen,bioinformatics,openvax
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: MacOS
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: beautifulsoup4>=4.12
Requires-Dist: datacache>=1.9.1
Requires-Dist: pysam>=0.22
Requires-Dist: requests>=2.28
Provides-Extra: test
Requires-Dist: pytest>=8; extra == "test"
Requires-Dist: ruff>=0.9; extra == "test"
Requires-Dist: build>=1; extra == "test"
Provides-Extra: docs
Requires-Dist: mkdocs<2,>=1.6; extra == "docs"
Dynamic: license-file

# osteosarc

Python library and command-line tool for the public [osteosarc.com](https://osteosarc.com/data/)
dataset: one patient's osteosarcoma sequencing, variant calls, cancer vaccines and
clinical history. Find files, pick variants and vaccine peptides, and fetch the
reads around a variant without downloading a whole BAM.

[Documentation](https://iskandr.github.io/osteosarc/) ·
[Key concepts](https://iskandr.github.io/osteosarc/concepts/) ·
[Command line](https://iskandr.github.io/osteosarc/cli/) ·
[Python API](https://iskandr.github.io/osteosarc/api/) ·
[Changelog](https://iskandr.github.io/osteosarc/changelog/)

## Features

- **Browse without downloading.** Search nearly 400,000 files by sample, timepoint and
  assay: bulk and single-cell RNA, exome, genome, Oxford Nanopore and PacBio.
- **Variants and vaccine peptides.** The site's variants with checked alleles, read
  counts, which pipelines found them, vaccine peptides and ELISPOT results.
- **Reads around a variant.** Copy just the reads you need out of a remote BAM into a
  small local one.
- **Corrected by default.** 35 fixes to known problems in the published data, each
  with its evidence, such as the MAP2 vaccine target's allele. Every load checks them against
  the snapshot's sources, and you can turn them off.
- **Clinical timeline.** Treatments, procedures, imaging, MRD and lab results as a
  text chart, or in an interactive terminal explorer.
- **Reproducible.** The site's metadata is saved as dated snapshots that reopen
  offline. The website changes; your results don't, until you sync again.
- **OpenVax integration.** Works with Varcode, Isovar, Topiary and Vaxrank, builds
  reproducible test BAMs, and includes a list of 637 candidate structural variants.

## Install

```sh
python -m pip install osteosarc
```

Needs Python 3.9+ on Linux or macOS, and [SAMtools](https://www.htslib.org/) on
your PATH to fetch reads.

## Quickstart

```python
from osteosarc import Dataset

data = Dataset.sync()  # Save today's website metadata (about 57 MB)
print(data.describe_samples())

rna = data.assets_for_sample("T0_tumor", kind="alignment", assay="rna-seq")
targets = data.variants(gene="DYNC1H1", status="ready")

source = rna["rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam"]
reads = data.extract_reads(source, variants=targets, padding=100)
print(reads.path)  # Local indexed BAM
```

The BAM key is the file's path in the dataset's public S3 bucket. Only the reads
near the variants are downloaded, into a small indexed BAM in your local cache.

Later, `Dataset.open()` reopens your most recent snapshot without a network
connection. The website changes over time, so snapshots are saved by download date:
`osteosarc snapshots` lists them, and `Dataset.open(date="2026-09")` or
`--snapshot 2026-09` picks the newest from that month.
[Get started](https://iskandr.github.io/osteosarc/#get-started) explains each step.

The same workflow from the terminal:

```sh
osteosarc sync
osteosarc samples
osteosarc assets --sample T0_tumor --kind alignment --assay rna-seq
osteosarc variants --gene DYNC1H1 --status ready
osteosarc reads rna-seq/reprocessed/BG003082/BG003082.Aligned.sortedByCoord.out.md.bam --variant DYNC1H1-chr14-101980529 --variant DYNC1H1-chr14-102030200 --padding 100
osteosarc explore
```

`osteosarc explore` opens an interactive browser for specimens, files, variants and
the timeline. Type `help` for commands and `quit` to leave.

## Guides

| I want to… | Read |
| --- | --- |
| Understand sample IDs, variant status, coordinates and corrections | [Key concepts](https://iskandr.github.io/osteosarc/concepts/) |
| Find RNA, DNA, single-cell or long-read files and read tables | [Find samples and files](https://iskandr.github.io/osteosarc/explore/) |
| Get alleles, read counts or vaccine peptides | [Select variants](https://iskandr.github.io/osteosarc/variants/) |
| Fetch, filter or pair reads by variant or region | [Extract reads](https://iskandr.github.io/osteosarc/reads/) |
| Browse treatments, specimens, MRD and labs | [Browse the timeline](https://iskandr.github.io/osteosarc/timeline/) |
| Pass data to Varcode, Isovar, Topiary or Vaxrank | [Use other libraries](https://iskandr.github.io/osteosarc/consumers/) |
| Build small, verifiable test BAMs | [Read fixtures](https://iskandr.github.io/osteosarc/fixtures/) |
| Explore candidate structural variants | [SV interest catalogue](https://iskandr.github.io/osteosarc/sv-interest/) |

## Data, license and citation

Osteosarc applies [source corrections](https://iskandr.github.io/osteosarc/curation/)
by default. Use `Dataset.open(corrections=False)` or
`osteosarc --no-corrections` to see the published values.

Code is Apache-2.0. The dataset is listed as CC0-1.0 in the
[AWS Open Data Registry](https://registry.opendata.aws/sid-osteosarc/).
Cite the dataset and your access date when using it.

## Development

```sh
python -m pip install -e '.[test]'
ruff check osteosarc tests scripts
python -m pytest -q
```

See [testing](https://iskandr.github.io/osteosarc/validation/) for documentation
builds and live-example checks.
