Metadata-Version: 2.5
Name: altar-alphamissense
Version: 0.1.0
Summary: AlphaMissense annotation-source binding for Altar
Project-URL: Documentation, https://kundajelab.github.io/altar/
Project-URL: Issues, https://github.com/kundajelab/altar/issues
Project-URL: Repository, https://github.com/kundajelab/altar
Author: Riya Sinha
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Typing :: Typed
Requires-Python: >=3.12
Requires-Dist: altar<0.2,>=0.1
Provides-Extra: test
Requires-Dist: pytest-asyncio>=0.24; extra == 'test'
Requires-Dist: pytest>=8; extra == 'test'
Description-Content-Type: text/markdown

# AlphaMissense binding

`altar-alphamissense` exposes precomputed AlphaMissense evidence through Altar's
`AnnotationSource` contract. AlphaMissense is a lookup source here, not a model
execution plugin: the published missense predictions already exist before an
Altar analysis begins.

The binding separates scientific normalization from storage. `AlphaMissenseSource`
aggregates transcript records into the stable per-variant columns consumed by
materialization, while an injected `AllelicRecordBackend` retrieves projected
native rows. Altar's generic `ParquetAllelicRecordBackend` is shared with
SpliceAI and other allelic sources; it contains no AlphaMissense logic.

```python
from altar.sources import AllelicKeyColumns, ParquetAllelicRecordBackend
from altar_alphamissense import AlphaMissenseRelease, AlphaMissenseSource

backend = ParquetAllelicRecordBackend(
    "./alphamissense",
    key_columns=AllelicKeyColumns("CHROM", "POS", "REF", "ALT"),
)
release = AlphaMissenseRelease(genome_build="hg38", source_release="zenodo-v1")
source = AlphaMissenseSource(backend, release=release, genome_build="hg38")
annotations = await source.annotate(["chr1:123:A:G"])
```

## Genome build and release

DeepMind publishes separate hg19 and hg38 files. `AlphaMissenseRelease` records
which one a backend serves and the exact upstream release. `genome_build` on the
source is the assembly of the variant IDs you will look up. The constructor
raises `ValueError` if it differs from the release's build.

The source also reads the native `genome` column of every record. If a record's
build differs from the release, `annotate` raises `ValueError` rather than
returning scores for the wrong assembly. A table that holds both builds must be
filtered to one, for example with `filters={"genome": "hg38"}` on the BigQuery
backend.

Every annotation carries `am_genome_build` and `am_source_release`, so
persisted AlphaMissense columns keep the build and release they came from.

For AlphaMissense records stored in BigQuery, use Altar's generic
`BigQueryAllelicRecordBackend`. The table needs the four key columns plus
`genome`, `uniprot_id`, `transcript_id`, `protein_variant`, `am_pathogenicity`,
and `am_class`. Its chromosome column must use the `chr`-prefixed spelling the
binding requests, such as `chr1`. A table spelled `1` makes staging raise
`ValueError`.

To include AlphaMissense in `BigQueryScoreStore` materialization and exports,
use `staged_annotation_source`. It runs the binding for the job's variants that
have AlphaMissense records and yields a source the store joins in SQL:

```python
from altar.sources import AllelicKeyColumns, BigQueryAllelicRecordBackend, staged_annotation_source
from altar_alphamissense import AlphaMissenseRelease, AlphaMissenseSource

backend = BigQueryAllelicRecordBackend(
    client,
    "project.reference.alphamissense",
    key_columns=AllelicKeyColumns("chr", "pos", "ref", "alt"),
    filters={"genome": "hg38"},
)
release = AlphaMissenseRelease(genome_build="hg38", source_release="zenodo-v1")
async with staged_annotation_source(
    AlphaMissenseSource(backend, release=release, genome_build="hg38"),
    client,
    staging_dataset="project.scratch",
    variants_relation="`project.scratch.job_variants`",
    coverage=backend,
) as source:
    ...  # pass `source` to the store's materialize and export calls
```

`ParquetAlphaMissenseBackend` remains as a thin compatibility preset for one
release window; new integrations should use the shared adapter directly.

Build that dataset from DeepMind's `AlphaMissense_hg38.tsv.gz` release with:

```console
altar-alphamissense-build --input AlphaMissense_hg38.tsv.gz --output ./alphamissense
```

`am_transcript_scores` is a typed, repeated struct that preserves every
matching transcript in descending pathogenicity order, with ties ordered by
transcript, UniProt accession, and protein variant. Each record has
`uniprot_id`, `transcript_id`, `protein_variant`, `am_pathogenicity`, and
`am_class`. Stores with native nested types keep it as an array of records;
SQLite stores it as JSON text without changing the logical schema.
`am_pathogenicity` and `am_class` are the summary from the most pathogenic
transcript and remain convenient scalar columns for filtering and
prioritization.
