Metadata-Version: 2.5
Name: altar-identity
Version: 0.1.0
Summary: Dependency-free variant and content identity shared by Altar and its model runtimes
Project-URL: Documentation, https://kundajelab.github.io/altar/
Project-URL: Issues, https://github.com/kundajelab/altar/issues
Project-URL: Repository, https://github.com/kundajelab/altar
Author: Riya Sinha
License-Expression: MIT
License-File: LICENSE
Keywords: bioinformatics,genomics,variant-identity
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Typing :: Typed
Requires-Python: >=3.9
Description-Content-Type: text/markdown

# altar-identity

`altar-identity` holds the identity rules that Altar core and every Altar model runtime must agree on. It uses
only the Python standard library and supports Python 3.9 and later, so a runtime pinned to an older
TensorFlow or PyTorch stack runs the same code as Altar core instead of keeping a copy.

Most users do not install it directly: `altar` depends on it and re-exports the variant names from
`altar.variants`, `altar.models`, and `altar.sources`.

## Variant identity

```python
from altar_identity import VariantKey, canonical_chromosome, canonical_variant_id

canonical_chromosome("MT")  # "chrM"
canonical_variant_id("1", 10, "a", "t")  # "chr1:10:A:T"
VariantKey.require_canonical("chr1:10:A:T")  # rejects aliases such as "1:10:A:T"
```

A key is `chromosome:position:REF:ALT` with a one-based position, and every field is ASCII. Surrounding ASCII
whitespace is trimmed from each field; other whitespace and non-ASCII text raise `VariantIdentityError`, a
`ValueError`, rather than being folded (an Arabic-Indic or full-width digit never becomes `1`).

- **Chromosome.** The `chr` prefix is optional and case-insensitive. Primary chromosomes are normalized in any
  case (`1`, `chr01`, `CHR1` → `chr1`; `M`, `MT` → `chrM`). Every other contig keeps its exact spelling after
  the prefix (`CHRUn_KI270302v1` → `chrUn_KI270302v1`), because reference contig names are case-sensitive and
  the key must match the name in the FASTA. The name uses the VCF contig-name characters: letters, digits, and
  `!#$%&*+./;=?@^_|~-`, which exclude `:`.
- **Position.** A positive `int`. In text (`VariantKey.parse`, variant files, and `parse_position` for other
  readers) it is ASCII digits `[0-9]+`: `+5`, `1_000`, `1e3`, and non-ASCII digits are rejected, although
  Python's `int()` accepts some of them.
- **Alleles.** Uppercased, then `[ACGTN]+`, the VCF base alphabet for REF and a concrete ALT. Symbolic
  (`<DEL>`), `*`, `.`, breakend, IUPAC-ambiguity and `-` alleles are rejected, and REF must differ from ALT.
  `N` is allowed in the key; reference validation decides whether a keyed variant can be scored.

The key does not left-align or trim indels, and it does not carry the genome build.

## Variant files

```python
from altar_identity import batched, read_variants

for batch in batched(read_variants("variants.tsv"), 1024):
    ...
```

Altar writes the variants for a model container as a headerless, tab-separated UTF-8 file with the columns
`chr`, `pos`, `ref`, `alt`, and `variant_id`. `read_variants` also accepts four columns, or any label in the
fifth, and never uses the fifth column. Fields are never quoted: a `"` is an ordinary character, and a field
cannot hold a tab or line break. A leading UTF-8 byte-order mark is ignored. The reader yields canonical
`VariantKey` values, skips blank lines, and raises `VariantFileError` naming the file and line for a malformed
row. By default a repeated `variant_id` is an
error; pass `duplicates="skip"` to keep the first occurrence or `duplicates="allow"` to yield every row.
`read_variant_rows` yields `VariantRow(line, key)` records instead, for callers that apply their own policy
(for example, SNVs only) and need to report the offending line. An empty or blank-only file yields nothing.

## Content identity

```python
from altar_identity import sha256_file, verify_file

digest = sha256_file("weights.h5")  # "sha256:<64 lowercase hex>"
verify_file("weights.h5", digest, label="weights")
```

A digest is `sha256:` followed by 64 lowercase hexadecimal digits and names the bytes of one regular file.
`verify_file` raises `DigestMismatchError` when the bytes differ, and also when the path is missing or is a
directory. Symlinks are followed. `parse_sha256_digest` and `is_sha256_digest` check the exact spelling, and
`SHA256_DIGEST_PATTERN` is the same rule as an unanchored regular expression for schemas that embed it.

A file that many tasks read, such as a reference genome on a shared volume, need not be hashed by every task.
`verify_file(..., trust_record=True, record=True)` accepts a file whose verification record is current and
writes a record after a successful hash. The record is a one-line file at `<path>.sha256-verified`:

```text
verified-file/1 sha256:<64 hex> size=<bytes> mtime=<seconds> ctime=<seconds> inode=<number>
```

Rewriting, replacing, truncating, or appending to the file changes a recorded value and invalidates the record.
`has_verification_record`, `write_verification_record`, and `verification_record_line` are the primitives Altar
core's `verify_file_digest`, its staging backends, and the runtimes share. Anyone who can write the storage
can also forge a record, so do not trust records where bytes first arrive, such as a download.

## Development

```bash
uv run --isolated --no-project --python 3.9 --with pytest --with-editable identity \
    python -m pytest identity/tests
```
