Metadata-Version: 2.4
Name: hitlist
Version: 1.65.0
Summary: Curated mass spectrometry evidence for MHC ligand data from IEDB and CEDAR
Author-email: Alex Rubinsteyn <alex.rubinsteyn@unc.edu>
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/pirl-unc/hitlist
Project-URL: Repository, https://github.com/pirl-unc/hitlist
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas
Requires-Dist: pyarrow
Requires-Dist: numpy>=1.22
Requires-Dist: PyYAML
Requires-Dist: tqdm
Requires-Dist: datacache>=1.15.0
Requires-Dist: pyensembl
Requires-Dist: mhcgnomes>=3.64.4
Requires-Dist: oncoref>=1.8.207
Provides-Extra: cta
Provides-Extra: alleles
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: pytest-cov; extra == "dev"
Requires-Dist: pytest-xdist; extra == "dev"
Requires-Dist: ruff<0.17,>=0.16; extra == "dev"
Dynamic: license-file

# hitlist

[![Tests](https://github.com/pirl-unc/hitlist/actions/workflows/tests.yml/badge.svg)](https://github.com/pirl-unc/hitlist/actions/workflows/tests.yml)
[![PyPI](https://img.shields.io/pypi/v/hitlist.svg)](https://pypi.org/project/hitlist/)

A curated, harmonized, **ML-training-ready** MHC ligand mass-spectrometry dataset.

hitlist ingests immunopeptidome data from [IEDB](https://www.iedb.org/), [CEDAR](https://cedar.iedb.org/), and paper supplementary tables (PRIDE/jPOSTrepo); partitions MS-eluted observations from in-vitro binding-assay measurements into two separate parquet files (so downstream consumers never silently conflate them); joins every MS observation to expert-curated sample metadata (HLA genotype, tissue, disease, perturbation, instrument); and ships both indexes as parquet + a pandas-friendly Python API.

## What's in the two indexes

After `hitlist build observations` (snapshot of the shipping 1.10.x default build):

### `observations.parquet` — MS-eluted immunopeptidome

| | |
|---|---|
| **Total observations** (MS-eluted, all species) | **4,053,693** |
| **Unique peptides** | **1,285,987** |
| Unique MHC alleles | 691 |
| MHC species covered | 21 |
| IEDB rows | 3,986,991 |
| CEDAR rows | 595 |
| Supplementary rows | 66,107 |

### `binding.parquet` — in-vitro binding-assay measurements

| | |
|---|---|
| **Total binding rows** (peptide microarray, refolding, MEDi, qualitative-tier) | **895,785** |
| **Unique peptides** | **258,199** |

The two indexes share the schema (including gene annotations from the peptide-mappings sidecar), but supplementary curation is MS-only — binding is pure IEDB/CEDAR.

### Human MHC-I breakdown

| | |
|---|---|
| Observations | 2,672,046 |
| Unique peptides | 748,386 |
| Mono-allelic (exact allele) | 579,096 obs / 300K peptides / 119 alleles |
| Multi-allelic with allele match | 450,399 obs |
| Multi-allelic with class-pool (`N of M alleles`) | 784,370 obs |
| Allele-resolved `sample_mhc` coverage | **74.8%** |

## What's curated

Those numbers rest on per-study human curation. Rather than trust the raw
IEDB/CEDAR `cell_name`, `disease`, and `tissue` fields — which carry mislabels,
free-text placeholders, and `Other`/`Unknown` sentinels — hitlist annotates each
study (keyed by PMID) in `pmid_overrides.yaml`. Known mislabels are corrected,
and every MS sample is tagged with the metadata a training pipeline actually
needs (HLA genotype, tissue, disease, perturbation, instrument) plus a reference
proteome for flanking and source-protein attribution. A few papers whose peptides
never reached IEDB are ingested directly from their PRIDE / jPOSTrepo
supplementary tables.

| What's curated | Count | Detail |
|---|---|---|
| Curated PMIDs (`pmid_overrides.yaml`) | **215** | Cover **96.0%** of all observations |
| ↳ `ms_samples` with per-sample metadata | 761 | HLA genotype, tissue, disease, instrument, plus the flat condition columns — knockout / cytokine / drug / infection / control role as scalars, not prose |
| ↳ `ms_samples` with 4-digit HLA typing | 579 | The subset with allele-level genotype |
| Supplementary CSVs ingested (PRIDE / jPOSTrepo) | 9 | 3 papers — listed below |
| Species reference proteomes | 22 | Ensembl ×4, UniProt ×18 |
| Viral reference proteomes | 33 viruses | 58 name aliases |

**Supplementary papers ingested** (data not in IEDB, read straight from
PRIDE/jPOSTrepo — 9 CSVs total):

- **Abelin 2019** — MAPTAC mono-allelic class I, class II, and DR-tissue (3)
- **Gómez-Zepeda 2024** — JY, HeLa, Raji, SK-MEL-37, and plasma (5)
- **Stražar 2023** — HLA-II (1)

## Install

```bash
pip install hitlist
```

## Quick start for ML training

```bash
# One-time: register IEDB + CEDAR downloads and build
hitlist data register iedb /path/to/mhc_ligand_full.csv
hitlist data register cedar /path/to/cedar-mhc-ligand-full.csv
hitlist build observations                           # a few minutes end-to-end;
                                             # writes observations.parquet +
                                             # binding.parquet + peptide_mappings.parquet

# Export training-ready CSVs
hitlist export training --include-evidence ms --class I --species "Homo sapiens" --mono-allelic \
    --min-allele-resolution four_digit -o mono_allelic_classI.csv

hitlist export training --include-evidence ms --class II --species "Homo sapiens" \
    -o classII_training.csv

# Presto-style flank-aware export: one row per (evidence row, peptide mapping)
hitlist export training --include-evidence both --class I --species "Homo sapiens" \
    --explode-mappings -o presto_training.parquet
```

`hitlist export training` does **not** create a new canonical store. It composes the existing `observations.parquet`, `binding.parquet`, and `peptide_mappings.parquet` indexes into one training-facing export surface. The low-level indexes keep their semantic boundaries; the training export gives downstream consumers one obvious API/CLI path when they want model-ready tables.

## Python API

`hitlist build observations` checks the source files and the curation inputs that define stored
annotations, including study overrides, tissue/cell-line metadata, and peptide-attribution CSVs.
Editing these inputs triggers a rebuild without `--force`. Upgrading from an older cache that
lacks curation fingerprints also triggers one rebuild.

### Training-data export

```python
from hitlist.export import generate_ms_observations_table
from hitlist.export import generate_training_table

# Mono-allelic human class I MS observations with ground-truth allele
mono_ms = generate_ms_observations_table(
    mhc_class="I",
    species="Homo sapiens",
    is_mono_allelic=True,
    min_allele_resolution="four_digit",
)

# Presto-style mapping-aware export: one row per (evidence row, peptide mapping)
presto = generate_training_table(
    include_evidence="both",
    mhc_class="I",
    species="Homo sapiens",
    map_source_proteins=True,
)
# columns now include: evidence_kind, evidence_row_id, protein_id, position,
# n_flank, c_flank, proteome, proteome_source
```

`generate_observations_table()` remains available as a backward-compatible alias.

Training exports preserve `evidence_row_id`, `evidence_source_id` and `evidence_kind`
even with a narrow `columns=` projection. Group mapping alternatives by
`evidence_row_id` to avoid counting protein mappings as additional observations.
Since 1.63.18, explicitly attributed donor/sample observations sharing an assay
have distinct, readable IDs, for example:
`ms:attributed:v1:http://www.iedb.org/assay/7578387|arm:31844290:mel3_13240_006`.
The compound key uses the original
source identifier and persistent curated `(pmid, condition_id)` when the original
attributed label uniquely identifies an arm; otherwise it uses the original
label scoped to its PMID (`|label:<pmid>:<escaped-label>`). Source locators and
attribution values are percent-escaped to avoid delimiter collisions.
Renaming a label while retaining its curated arm ID
preserves the ID. Resolving a previously unmatched label to an arm changes its ID.
IDs do not depend on row order, current filters or mapping alternatives.

`evidence_source_id` retains the previous common `<kind>:<assay_iri>` identifier
for migration and source tracing. Rows without explicit attribution also retain
that value as `evidence_row_id`. Older indexes lacking assay IRIs still fall back
to reference IRIs or row positions; those legacy fallbacks cannot guarantee
unique, filter-stable observation IDs. An observation/curated arm ID is not a
verified biological specimen ID or proof of experimental independence.

Since 1.64.0, rebuilt indexes retain every source contributor, and training
exports carry `provenance_id` and reviewed experimental lineage. Create a
versioned bundle with `hitlist export training --bundle DIR`, then compare
partitions with `hitlist audit-splits`. Unknown specimen/donor identities remain
explicit, and incomplete coverage cannot certify independence. See
[Training provenance and split audits](docs/training-provenance.md) for the
manifest contract, policies and legacy-index limitations (#622, #616).

Gene queries can mix symbols and Ensembl IDs: `gene=["PRAME", "ENSG00000198681"]` selects
peptides matching either gene. An additional `peptide=` filter narrows that union. Mapping-expanded
training exports retain only mappings for the selected genes, even when a peptide also maps to
an unrelated gene. Separately supplied `gene_name` and `gene_id` filters on the low-level loaders
remain conjunctive.

Allele-set filters use the same normalization in raw loaders and exports: `mhc_allele_in_set="A*02:01"`
and `mhc_allele_in_set="HLA-A*02:01"` select the same evidence. Explicit empty allele-set queries
raise `ValueError` consistently.

Species filters accept any variant — `"Homo sapiens"`, `"human"`, `"homo_sapiens"`, `"Homo sapiens (human)"` all work.

MS, binding, training, and peptide-summary exports support independent `source_species` and
`host_species` filters, plus `exclude_chimeric=True`. The existing `species` filter selects MHC
species. For example, `generate_training_table(species="human", source_species="mouse")` selects
mouse-source peptides presented on human MHC. The matching CLI flags are `--source-species`,
`--host-species`, and `--exclude-chimeric`; quote multiword species names. Defaults retain all
species and chimeric systems. Source filtering and system flags both fall back to the raw
`species` column when `source_organism` is blank; raw source annotations remain available.

### Raw observations loading

```python
from hitlist.observations import (
    load_ms_observations,     # MS-eluted immunopeptidome
    load_binding,             # in-vitro binding-assay measurements
    load_all_evidence,        # union, tagged with an evidence_kind column
    is_built, is_binding_built,
    observations_cache_is_current,  # current vs. legacy; never builds or prints
    observations_path, binding_path,
)

# MS-elution (the default training-data path)
df = load_ms_observations()                       # everything (MS-eluted only)
df = load_ms_observations(mhc_class="I")          # class I only
df = load_ms_observations(species="Homo sapiens") # human only
df = load_ms_observations(source="iedb")          # filter by source
df = load_ms_observations(columns=["peptide", "mhc_restriction", "src_cancer"])

# Binding assays — same filter API, reads binding.parquet
bd = load_binding(mhc_class="I", mhc_restriction="HLA-A*02:01")

# Union — for affinity-predictor training, or UI flags that want both.
# Rows are tagged with evidence_kind ∈ {"ms", "binding"}.
both = load_all_evidence(gene_name="PRAME", mhc_class="I")
both["evidence_kind"].value_counts()
```

`load_observations()` remains available as a backward-compatible alias.

`is_built()` only says a parquet exists. To learn whether it is *current* — the same
verdict `build_observations(force=False)` reaches before skipping — call
`observations_cache_is_current()`. It returns `True`, `False`, or `None` when no IEDB/CEDAR
source is registered and validity cannot be judged. `hitlist.mappings.mappings_cache_is_current()`
answers the same for the `peptide_mappings.parquet` sidecar. Neither prints, writes, or
downloads, so a consumer can announce a rebuild before starting one instead of buffering
the builder's output and inferring afterwards.

### Building / curation

```python
from hitlist.builder import build_observations
from hitlist.curation import (
    classify_ms_row,
    normalize_species,
    normalize_allele,
    load_pmid_overrides,
)
from hitlist.supplement import scan_supplementary, load_supplementary_manifest

build_observations(with_flanking=True, use_uniprot_search=True, force=False)
normalize_species("human")           # → "Homo sapiens"
normalize_allele("H-2Kb")            # → "H2-K*b"
scan_supplementary()                 # DataFrame of curated paper-supplement peptides
```

### Peptide → protein attribution and flanking context

`hitlist build observations` always produces three parquet files (use `--no-mappings` to skip `peptide_mappings.parquet`):

- `observations.parquet` — one row per assay observation
- `binding.parquet` — one row per binding-assay observation
- `peptide_mappings.parquet` — one row per (peptide, protein, position)

They land in the data directory (`hitlist data dirs`; see
[Where hitlist keeps data](#where-hitlist-keeps-data)).

The mappings sidecar **preserves multi-mapping** so a peptide shared by MAGEA1/A4/A10/A12 keeps every paralog. Ensembl mappings include `gene_biotype`: the default index covers ordinary `protein_coding` genes plus the coding `IG_V/D/J/C_gene` and `TR_V/D/J/C_gene` biotypes, while excluding pseudogenes. These receptor records are germline segments; Ensembl does not contain a donor's recombined receptor, so peptides spanning a V(D)J junction cannot map. Observations additionally carry semicolon-joined identity columns:

| column | example |
|---|---|
| `gene_names` | `MAGEA4;MAGEA10` |
| `gene_ids` | `ENSG00000147381;ENSG00000124260` |
| `protein_ids` | `P43359;P43363` |
| `n_source_proteins` | `2` |

```python
from hitlist.observations import load_ms_observations
from hitlist.mappings import load_peptide_mappings

# Central columns — fast for everyday filters (uses mappings sidecar for pushdown)
df = load_ms_observations(gene_name="PRAME")

# Long form for paralog / position / flank analysis
mappings = load_peptide_mappings(gene_name="MAGEA4")

# Receptor-derived mappings remain distinguishable from ordinary proteins
receptor_mappings = load_peptide_mappings(
    gene_biotype=["IG_V_gene", "IG_D_gene", "IG_J_gene", "IG_C_gene",
                  "TR_V_gene", "TR_D_gene", "TR_J_gene", "TR_C_gene"]
)
# columns include: peptide, protein_id, gene_name, gene_id, gene_biotype,
# transcript_id, position, n_flank, c_flank, proteome
```

For ad-hoc queries without building the full table:

```python
from hitlist.proteome import ProteomeIndex

idx = ProteomeIndex.from_ensembl(release=112, species="human")  # coding + germline IG/TR
flanking = idx.map_peptides(["SLLMWITQC", "GILGFVFTL"], flank=10)

# Explicit compatibility mode for the historical protein-coding-only index:
ordinary_only = ProteomeIndex.from_ensembl(release=112, biotype="protein_coding")
```

### Proteome registry / UniProt resolution

```python
from hitlist.downloads import (
    lookup_proteome,           # org string → registry entry (dict)
    fetch_species_proteome,    # download FASTA and cache to <data dir>/proteomes/
    resolve_proteome_via_uniprot,  # direct UniProt REST lookup
    list_proteomes,            # manifest section
)

lookup_proteome("Mycobacterium tuberculosis")
# → {'kind': 'uniprot', 'proteome_id': 'UP000001584', ...}  # H37Rv reference
```

## Output schema — `generate_ms_observations_table()`

| Column | Meaning |
|---|---|
| `peptide` | Amino acid sequence |
| `mhc_restriction` | Allele from IEDB (may be `"HLA class I"` for multi-allelic studies) |
| `sample_mhc` | Experimental MHC candidates; may be a selected restriction or an unresolved study/class pool, not a complete genotype |
| `sample_mhc_origin` | `sample` (this arm's own candidates), `class_pool` (the study's class-wide union), blank (none reached the row), or `not_applicable` on a binding row in the training table. The fact `sample_match_type` cannot give you: that column says whether the study *has* a pool, not whether this row took it |
| `mhc_basis` | Source-reviewed `selected_restriction` or `sample_typing`. Blank on an *exported row* means either unreviewed or that the claim was dropped because `sample_mhc` is not that arm's — `sample_mhc_origin` tells you which |
| `mhc_genotype`, `mhc_genotype_cell`, `mhc_genotype_source` | Independently sourced cellular typing, its cell/donor identity and citation; never pooled across cells |
| `mhc_genotype_reported_loci`, `mhc_genotype_complete_loci` | Loci with molecular typing versus explicitly complete typing; missing loci remain unknown |
| `mhc_class` | Canonical `I`, `II`, or `non-classical`; molecule-derived when possible |
| `mhc_class_reported` | Source-reported class, retained verbatim for auditability |
| `mhc_class_source`, `mhc_class_corrected` | Whether class came from one molecule, a consistent donor set, or source fallback; whether it corrected the source |
| `mhc_species` | Canonical MHC species (mhcgnomes plus explicit per-study context for ambiguous names) |
| `mhc_species_source`, `mhc_species_context_disagrees` | Species-resolution provenance and explicit context-conflict signal |
| `restriction_evidence` | How the named peptide-to-MHC restriction was established: `experimental`, `monoallelic`, `predicted`, or `unknown` |
| `serotype`, `serotypes` | Canonical serotype and full membership, preserving the MHC species prefix (for example `HLA-A2` or `BoLA-A18`). An allele can carry a locus serotype and a public epitope such as `Bw4`. |
| `serotype_source` | Whether the serotype is primary data or a projection: `reported` (the study typed serologically; no molecule was measured, which is why the allele fields are empty), `derived` (computed from a named molecule through mhcgnomes' membership table, so only as current as the installed mhcgnomes — the build records `mhcgnomes_version` in `observations_meta.json`), `donor_set` (a union over a donor's typed alleles, making the serotype a candidate rather than the restriction's identity), or empty (no serotype) |
| `is_monoallelic` | True if sample has a single transfected allele (721.221, C1R, K562, MAPTAC…) |
| `has_peptide_level_allele` | True if `mhc_restriction` is a specific allele (not `"HLA class I"`) |
| `is_potential_contaminant` | True for MS-eluted peptides that failed NetMHCpan binding prediction |
| `sample_match_type` | How `sample_mhc` was populated (see below) |
| `matched_sample_count` | Number of curated samples for this PMID |
| `src_cancer`, `src_healthy_tissue`, `src_ebv_lcl`, ... | Mutually-exclusive biological source categories |
| `source` | `iedb`, `cedar`, or `supplement` |
| `source_organism`, `reference_title`, `cell_name`, `source_tissue`, `disease` | IEDB sample context |
| `instrument`, `instrument_type`, `acquisition_mode`, `fragmentation`, `labeling`, `ip_antibody` | MS acquisition from `ms_samples` curation |
| `gene_names`, `gene_ids`, `protein_ids`, `n_source_proteins` | Multi-mapping peptide → source-protein attribution (always populated; use `peptide_mappings.parquet` for long-form positions + flanks) |

### `sample_match_type` — join provenance

Reported gene/locus restrictions such as HLA-DR, DQ and DP also retain a separate
species/class/locus-compatible `mhc_allele_set`, with explicit inference provenance.
See [allele candidate matching and coverage](docs/allele-candidates.md). These
candidates do not upgrade the reported restriction or establish a measured
peptide-to-allele assignment.

| Value | Meaning | Training-grade? |
|---|---|---|
| `allele_match` | IEDB recorded an allele matching the curated experimental candidates | Subject to the row's restriction evidence and sample attribution |
| `single_sample_fallback` | Study has exactly one eligible sample; its reported candidates supply `sample_mhc` | Candidate set for that sample, with source precision retained |
| `pmid_class_pool` | Class-based attribution; unresolved rows may carry the union of class-matching candidates | An unresolved pool must not be used as one sample's genotype |
| `unmatched` | No curated sample for this PMID, or all samples have `mhc: unknown` | No — `sample_mhc` empty |

`sample_attribution` says *how* the row reached its sample, and one value is
worth reading before trusting an arm. `elution_conditions_excluded` means the
row's own deposited elution statement named arms that exclude the one its
allele matched — for a pan-class-II elution of a line whose class-II arms are
all transduced, a peptide deposited under the untransduced statement. Those
rows carry no `sample_label` and take the class pool, because the arm they came
from is named by the deposit and absent from curation (#565, #567). Treat them
as unattributed, not as the arm their allele happens to fit.

## Biological source classification

Every observation is classified by mutually-exclusive biological source category:

| Category | Flag | Rule |
|---|---|---|
| Cancer | `src_cancer` | Tumor tissue, cancer patient biofluids, or non-EBV cell lines |
| Adjacent to tumor | `src_adjacent_to_tumor` | Surgically resected "normal" tissue (per-PMID override) |
| Activated APC | `src_activated_apc` | Monocyte-derived DCs/macrophages with pharmacological activation |
| Healthy somatic | `src_healthy_tissue` | Direct ex vivo, healthy donor, non-reproductive, non-thymic |
| Healthy thymus | `src_healthy_thymus` | Direct ex vivo thymus (expected for CTAs, AIRE-mediated) |
| Healthy reproductive | `src_healthy_reproductive` | Direct ex vivo testis, ovary (expected for CTAs) |
| EBV-LCL | `src_ebv_lcl` | EBV-transformed B-cell lines |
| Cell line | `src_cell_line` | Any cultured cell line |

**Cancer-specific** = `src_cancer AND NOT src_healthy_tissue`. Thymus, reproductive tissue, adjacent tissue, EBV-LCLs, and activated APCs do NOT disqualify a peptide from being cancer-specific.

## CLI reference

### Data management

```bash
hitlist data register <name> <path> [-d DESCRIPTION]    # register a local file
hitlist data fetch <name> [--force]                     # download a known dataset (IEDB/CEDAR/viral FASTAs)
hitlist data refresh <name>                             # re-download
hitlist data info <name>                                # detailed metadata (JSON)
hitlist data path <name>                                # print the registered path
hitlist data remove <name> [--delete]                   # unregister (optionally delete file)
hitlist data list                                       # registered datasets + mirrored assets
hitlist data list --all --json                          # every cache file and its provenance
hitlist data list --verify                              # verify trusted mirrored-asset hashes
hitlist data available                                  # show all known datasets
hitlist data dirs                                       # every directory hitlist uses, and why
```

### Where hitlist keeps data

`hitlist data dirs` is the complete answer; the short version:

See [downloads and cache inspection](docs/downloads.md) for resumability,
read-only inventory, integrity statuses, and the Python download adapter.

| what | where |
|---|---|
| built indexes (`observations`/`binding`/`bulk_proteomics`/`line_expression` parquets, `manifest.json`, downloaded proteomes) | the data directory — see the resolution order below |
| mirrored data assets (paper-derived CSVs) | `datacache`'s cache dir for the `hitlist` subdir; **not** moved by `HITLIST_DATA_DIR` |
| proteome index cache (`*.pkl`) | `<data directory>/proteome_index_cache`, unless explicitly set with `hitlist.proteome.set_disk_cache_dir()` |

The data directory resolves in this order:

1. `hitlist.downloads.set_data_dir()`
2. the `HITLIST_DATA_DIR` env var — the directory itself, exactly as it always
   has been; whitespace is stripped, `~` is expanded, and an empty value means
   unset (it used to resolve to the process's working directory)
3. an existing, **populated** `~/.hitlist` — the legacy location, still fully
   supported, so an install that already has a corpus there keeps using it (and
   says so once per process)
4. otherwise `datacache.get_data_dir(subdir="hitlist")` — `~/Library/Caches/hitlist`
   on macOS, `~/.cache/hitlist` on Linux, the convention pyensembl and the rest of
   the openvax ecosystem already use

A `~/.hitlist` counts as **populated** when it holds a regular file at the top
level, or a subdirectory holding a regular file. The test is structural on
purpose — a list of known artifact names would drift, and would already have
missed the user whose only data is `gene_cache/hgnc_lookups.json`.

Merely existing is not enough: older releases created `~/.hitlist` on *every*
call to the path helper, and `<data dir>/proteomes/` and `<data dir>/gene_cache/`
are still created eagerly, so plenty of installs have one holding nothing but
empty folders.

Unconfigured installs keep their existing populated legacy directory, including
one containing only proteome index cache files. The proteome cache now follows
`HITLIST_DATA_DIR` and `set_data_dir()`; an existing explicit data-dir setting
therefore also selects the proteome cache location. Old cache files remain in
place. Copy `proteome_index_cache/` into the selected data directory to reuse
them, or let the cache rebuild. `set_disk_cache_dir(None)` restores this default
and follows subsequent data-dir changes.

To move an existing corpus, set
`HITLIST_DATA_DIR` to the new location and copy the old directory's **entire**
contents there. Moving only the parquets leaves `manifest.json` behind, which
keeps rule 3 pointed at a `~/.hitlist` that no longer has any indexes in it.

### Build the observations table

```bash
hitlist build observations [--force]                            # ~90s full scan with tqdm progress
hitlist build observations                                      # always builds peptide_mappings.parquet
hitlist build observations --use-uniprot                        # broader proteome coverage via UniProt REST
hitlist build observations --no-mappings                        # skip mapping step (faster, no gene attribution)
hitlist build observations --no-fetch-proteomes                 # don't auto-download missing proteomes
hitlist build observations --proteome-release 112               # Ensembl release for human/mouse/rat
```

### Proteome management

```bash
hitlist data fetch-proteomes [--min-observations N] [--use-uniprot] [--force]
hitlist data list-proteomes
```

### Export

```bash
hitlist export ms [filters...] -o train.csv             # MS immunopeptidome + sample metadata
hitlist export ms -o train.parquet                      # parquet output supported
hitlist export peptide-summary --gene PRAME --serotype A24   # per-peptide support for one allele/serotype
hitlist export binding [filters...] -o binding.csv      # binding-assay index (separate from MS)
hitlist export training [filters...] -o training.csv    # unified training export from canonical indexes
hitlist export peptide-counts --by class                # species x class peptide counts
hitlist export peptide-counts --by study                # peptide counts per study/PMID
hitlist samples [--class I|II]                          # per-sample conditions (promoted from `export samples`)
hitlist qc normalization                                # validate YAML alleles with mhcgnomes
hitlist qc mhc-tokens                                   # unparseable MHC tokens across all modalities
hitlist qc resolution                                   # allele-resolution histogram for IEDB/CEDAR
```

`hitlist export ms` is the canonical name for the MS observations export; `hitlist
export observations` remains as a backward-compatible alias.

### Canonical indexes and the training export

Each `hitlist build observations` writes **three** parquet files to the data
directory (`hitlist data dirs`):

- `observations.parquet` — MS-eluted immunopeptidome (IEDB + CEDAR + curated supplementary).
- `binding.parquet` — binding-assay rows (peptide microarray, refolding, MEDi, and
  quantitative-tier measurements like `Positive-High/Intermediate/Low`).
- `peptide_mappings.parquet` — long-form peptide → protein/position/flank mappings.

The canonical indexes are never silently mixed. Supplementary data is MS-only.
Use `hitlist export observations` and `hitlist export binding` when you want
the raw evidence families separately. Use `hitlist export training` or
`generate_training_table(...)` when you want a composed model-facing export
with `evidence_kind` tagging and optional mapping explosion for flank-aware
training pipelines.

### Filters on `hitlist export observations`

| Flag | Values |
|---|---|
| `--class` | `I`, `II`, `non classical` |
| `--species` | Any species variant (normalized via mhcgnomes) |
| `--mono-allelic` / `--multi-allelic` | Filter on `is_monoallelic` |
| `--instrument-type` | `Orbitrap`, `timsTOF`, `TOF`, `QqQ`, ... |
| `--acquisition-mode` | `DDA`, `DIA`, `PRM` |
| `--min-allele-resolution` | `four_digit`, `two_digit`, `serological`, `class_only` |
| `--mhc-allele` | Exact match on `mhc_restriction` after allele normalization. Repeatable / comma-separated. |
| `--restriction-evidence` | Restriction evidence: `experimental`, `monoallelic`, `predicted`, or `unknown`. Independent of allele-set provenance. Repeatable. |
| `--serotype-source` | Whether the serotype is primary data or a projection: `reported` (serologically typed study), `derived` (computed from a named molecule), `donor_set` (union over a donor's typed alleles). Pair with `--serotype` to keep a serotype query on measured typings. Repeatable. |
| `--gene` | Symbol, Ensembl ID, or old alias (HGNC synonym lookup). Repeatable / comma-separated. Requires the mappings sidecar (default-on at build). |
| `--gene-name` | Exact match on `gene_name` column (no HGNC lookup) |
| `--gene-id` | Exact match on `gene_id` column (ENSG) |
| `--peptide` | Exact match on `peptide` sequence. Repeatable / comma-separated. |
| `--serotype` | MHC serotype: use a species prefix for non-human names (`BoLA-A18`, `Patr-DR1`); HLA names may omit it (`A24`, `B57`, `DR15`, `Bw4`, `Bw6`). Matches any serotype the allele belongs to, so `--serotype Bw4` returns A\*24:02, B\*27:05, B\*57:01, etc. HLA split serotypes are rolled into their broad parent, so `--serotype A24` also matches A\*24:03 (serotype `A2403`). Repeatable / comma-separated. |
| `--exclude-class-label-suspect` | Drop rows where the curated class disagrees with peptide length severely enough to be flagged `suspect` or `implausible` (`mhc_class_label_severity`). |
| `--exclude-class-label-implausible` | Strict-cleaning variant — drops only `implausible` rows (class-I ≥18aa or ≤7aa, class-II ≤4 or ≥45aa). Keeps borderline + suspect tiers, useful when bulged class-I 15-17aa peptides should be retained. |
| `--apm-only` | Filter to peptide rows from samples where any APM gene was perturbed (`apm_perturbed=True`). Reflects the sample's *own* condition; the parent study's perturbation panel is carried separately in `study_apm_perturbed` / `study_apm_genes`. |
| `--output` / `-o` | `.csv` or `.parquet` |

For filtering a frame already in memory, `hitlist.normalize_serotype_query()` uses
the same spelling rules as the loaders and exports: `"a2"` becomes `"HLA-A2"`,
and `"bola-a18"` becomes `"BoLA-A18"`.

All filters are pushed down to the parquet reader (pyarrow), so `--gene PRAME` reads
only the matching row groups — typically milliseconds rather than a full table scan.
Examples:
- `hitlist export ms --gene PRAME --class I -o prame_classI.csv`
- `hitlist export ms --gene "MART-1"` (HGNC resolves to `MLANA`)
- `hitlist export ms --mhc-allele HLA-A*02:01 --mono-allelic`
- `hitlist export ms --serotype A24` (locus-specific)
- `hitlist export ms --serotype Bw4` (public epitope — A*23/24/25/32, B*13/27/44/51/52/53/57/58)

### `hitlist export peptide-summary`

Collapses the MS observations to **one row per peptide**, scoring how strongly
each peptide is supported on a single target allele or serotype. Answers
questions like *"which PRAME peptides might be presented on A24 in cancers?"*

```bash
hitlist export peptide-summary --gene PRAME --serotype A24 -o prame_a24.csv
hitlist export peptide-summary --gene PRAME --mhc-allele HLA-A*24:02
```

Requires exactly one of `--mhc-allele` / `--serotype` (a single value) plus a
gene/peptide scope filter. Each peptide's support is split into ranked tiers,
strongest first:

| Bucket column | Meaning |
|---|---|
| `n_mono_exact_rows` | Mono-allelic elution on the exact target allele (strongest) |
| `n_multi_exact_rows` | Multi-allelic elution including the exact target allele |
| `n_mono_serotype_rows` / `n_multi_serotype_rows` | Elution on a same-serotype allele |
| `n_class_only_sample_allele_rows` | Class-only row whose experimental candidates include the target allele |
| `n_class_only_sample_serotype_rows` | Class-only row whose experimental candidates include a matching serotype |
| `n_unknown_allele_rows` | Class-only row with no usable reported candidates (weakest) |

`best_support` names the strongest tier with any rows. A `cancer` / `healthy` /
`adjacent` / `other` source breakdown (`n_cancer_rows`, ...) comes along for free,
so you can prioritize peptides seen in tumors over healthy tissue.

### Filters on `hitlist export binding`

Same shape as the observations filters minus the MS-specific ones
(`--mono-allelic`, `--instrument-type`, `--acquisition-mode`). `--source`
accepts only `iedb` or `cedar` — supplementary data is MS-only and never
appears in the binding index.

```bash
hitlist export binding --gene PRAME --class I -o prame_binding.csv
hitlist export binding --mhc-allele HLA-A*02:01 --serotype Bw4
```

### Filters on `hitlist export training`

`hitlist export training` exposes the shared pMHC filters plus two export-shape controls:

- `--include-evidence ms|binding|both` chooses which canonical evidence families to compose.
- `--explode-mappings` expands the output to one row per `(evidence row, peptide mapping)` with `protein_id`, `gene_biotype`, `position`, `n_flank`, `c_flank`, `proteome`, and `proteome_source`.

MS-specific filters (`--mono-allelic`, `--instrument-type`, `--acquisition-mode`) apply only to the MS slice. Binding rows never gain fake sample context; they remain tagged as `evidence_kind="binding"` with `sample_match_type="not_applicable"`.

Binding-specific filters (`--assay-method`, `--response-measured`,
`--measurement-units`, `--has-quantitative-value` / `--qualitative-only`, and
`--quantitative-value-min` / `--quantitative-value-max`) apply only to the binding
slice. With `--include-evidence both`, they retain the independently selected MS
rows; with `ms`, they have no effect. Methods use case-insensitive substring
matching; endpoints and units use case-insensitive exact matching. Multiple
values select any listed value; different filters combine with AND. Numeric
bounds are inclusive and exclude missing values. No units are converted and
inequality qualifiers are preserved. The same options work with `--bundle DIR`
and are recorded in its manifest.

```bash
hitlist export training --include-evidence both --gene PRAME --class I -o prame_training.csv
hitlist export training --include-evidence ms --mono-allelic --class I -o mono_ms.csv
hitlist export training --include-evidence both --explode-mappings -o presto_training.parquet
```

### Selecting MHC-I affinity measurements

Choose an explicit endpoint allowlist and units. This example selects human
MHC-I IC50 and KD measurements reported in nM; it deliberately excludes EC50,
KD proxies such as `dissociation constant KD (~IC50)`, broad `MHC binding`
labels, half-life, structure and thermal-stability endpoints. Inspect the
reported vocabulary and define the endpoint policy for your model. IC50 and KD
remain distinct endpoints even when selected together.

```python
from hitlist.export import generate_training_table

affinity = generate_training_table(
    include_evidence="binding",
    mhc_class="I",
    species="human",
    min_allele_resolution="four_digit",
    response_measured=[
        "half maximal inhibitory concentration (IC50)",
        "dissociation constant KD",
    ],
    measurement_units="nM",
    has_quantitative_value=True,
)

# Explicit policy for a model that accepts only exact, positive numeric targets.
# Preserve endpoint and original measurement columns alongside every target.
exact = affinity[
    affinity.measurement_inequality.eq("=") & affinity.quantitative_value.gt(0)
].copy()
censored = affinity[affinity.measurement_inequality.isin(["<", "<=", ">", ">="])].copy()
unresolved_qualifier = affinity[
    ~affinity.measurement_inequality.isin(["=", "<", "<=", ">", ">="])
].copy()
```

Equivalent selection before the exact/censored partition:

```bash
hitlist export training --include-evidence binding --class I --species human \
    --min-allele-resolution four_digit \
    --response-measured 'half maximal inhibitory concentration (IC50)' \
    --response-measured 'dissociation constant KD' \
    --measurement-units nM --has-quantitative-value -o affinity.parquet
```

`has_quantitative_value=True` alone includes any endpoint with a number, including
half-life or thermal stability. Unit selection alone also does not identify an
affinity endpoint. Numeric bounds filter the **reported number**: a `>5000` row
passes `quantitative_value_max=5000`, but its measurement remains greater than
5000. Use censored rows only with a model/loss that handles their bounds; do not
replace their qualifiers with equality. Blank or unfamiliar qualifiers need
review. `has_quantitative_value=False` (`--qualitative-only`) selects rows without
a numeric value, preserving `qualitative_measurement`; those labels need their
own explicit policy and are not invented numeric targets. The canonical indexes
retain all endpoint families. Column projection is optional; retain
`response_measured`, `measurement_units`, `quantitative_measurement`,
`quantitative_value`, `measurement_inequality`, and `qualitative_measurement`
when passing selected measurements downstream.

### Sample-level expression anchors (issue #140)

For line-like `ms_samples` (C1R, 721.221, JY and sibling EBV-LCLs, HAP1,
HeLa, HEK293, THP-1, SaOS-2, A375, K562, GM12878, and common engineered
derivatives) `hitlist.line_expression` resolves every sample to a stable
expression backend + key via a 6-tier fallback hierarchy:

1. **exact-line RNA / transcript quant** (registry hit with shipped data)
2. **parent-line / engineered-derivative RNA** (e.g. HeLa.ABC-KO → HeLa)
3. **line-family class anchor** (EBV-LCL → GM12878; mono-allelic host → K562)
4. **cancer-type surrogate** (caller-supplied `pirlygenes` backend)
5. **broad tissue / lineage surrogate** (HPA)
6. **`no_expression_anchor`**

Four provenance columns — `expression_backend`, `expression_key`,
`expression_match_tier`, `expression_parent_key` — are written on every
row so downstream tooling can distinguish "exact JY RNA" from "generic
EBV-LCL stand-in" from "melanoma cohort surrogate" instead of treating
them as equally trustworthy.

Tier 1 requires that the profiled material *is* the registered line, so an
engineered sample never resolves there. Engineering comes from curation, not
from parsing the label: any intervention in `condition_knockout_genes`,
`condition_knockdown_genes`, `condition_overexpression_genes`,
`condition_genetic_variants`, `condition_transfection` or
`condition_transduction`, or an introduced-MHC `condition_mhc_context`
(`monoallelic`, `mhc_transfectant`, `mhc_coexpression`). Such a sample
resolves at tier 2 against the same data even when its label matches the
reference line's alias — `HeLa-CIITA (Mock)`, `SaOS-2 + TP53 R175H`,
`K562 transfectant DPB1*01:01/DPA1*02:01` — with `expression_parent_key`
naming the material that supplied the RNA. `resolve_sample_expression_anchor`
derives this itself when given the `pmid` and `sample_label` of a curated arm;
pass `engineered=` to override. A row that names no arm — an observation
the export could not attribute to one of HAP1's wild-type and knockout arms,
say — is not known to be unmodified when the study has an engineered arm of
the *same reference line* as the row's anchor (lines are matched through the
registry, so aliases count), and also resolves at tier 2 with a reason saying
the arm is unresolved. An unattributed JY row keeps tier 1 in a study whose
only engineered arm is a Raji transfectant.

CLI:

```bash
hitlist export samples --with-expression-anchors -o sample_anchors.csv

hitlist export training \
    --include-evidence ms --class I --mono-allelic \
    --with-peptide-origin \
    -o training_with_origin.parquet

hitlist export line-expression --line-key GM12878 --gene-name TP53
```

`--with-peptide-origin` attaches, per row, the argmax-TPM candidate gene
(`peptide_origin_gene`, `peptide_origin_tpm`) for the peptide in the
sample's resolved line. When transcript-level TPM is available the
score is the sum across **only those transcripts whose translation
actually contains the peptide** — isoforms that splice out the peptide
contribute zero, and a transcript is counted once regardless of how many
times the peptide appears in its protein.

Fetch and build the optional DepMap 24Q4 gene/transcript expression bundle
(about 4.7 GB). It provides RNA anchors for HeLa, A375, SaOS-2, THP-1 and
K562, and HAP1 transcript rows. This release has no HEK293 model entry.
HAP1's DepMap gene row ships with the package, so HAP1 and its engineered
KO panel resolve without the bundle (tier 1 and tier 2 respectively).

```bash
hitlist data fetch depmap
```

The bundle includes the model, profile and default-profile mappings needed
for transcript rows. Existing registered copies are reused. Individual
files can still be fetched or registered under `depmap_rna`,
`depmap_rna_transcript`, `depmap_models`, `depmap_profiles` and
`depmap_default_profiles`. RNA anchors report only sources with available
rows; missing optional data falls back to the next applicable tier.

`line_expression.parquet` is stamped with a content fingerprint of every
packaged build input — the anchor registry, `sources.yaml`, the packaged CSVs
and the ENSG → HGNC symbol map — plus an index build version. Reads compare
the stamp with the installed release and never rewrite the file:

- **current**: the index is used as built.
- **no stamp** (every index written before this release) **or same build
  version with different packaged inputs**: reads use this release's packaged
  sources plus the index's downloaded (e.g. DepMap) rows, re-enriched, for
  the (line, source) pairs the current registry still lists and whose
  provenance meets the current builder's invariants. An upgrade or an anchors
  edit therefore does not cost a DepMap user HeLa or K562, while rows the
  registry no longer lists (HAP1 under `DepMap_24Q4_gene`) are left out, as
  are rows from before DepMap default-profile selection (#357): an index with
  no `profile_id` column, or a transcript row without its `PR-…` profile.
- **another build version, or an unreadable stamp**: packaged sources only.

Without a usable index, the packaged sources are served in the same row shape
a build writes (source metadata, `parent_line_key`, and gene symbols from the
packaged map), so `gene_name` filters reach them. The first read of a
non-current index warns once, naming what was left out and the cheapest
rebuild: `hitlist data fetch depmap` when the index held downloaded rows and
every DepMap file is still registered (a purely local rebuild); otherwise a
local `hitlist.builder.build_line_expression()`, with the exact DepMap files
a fetch would still download. `line_expression_index_is_current()` checks
without warning; `is_line_expression_built()` only says a file exists.

### A note on mono-allelic curation

Mono-allelic is a **PMID-level** flag, not a per-sample property. Curation
lives in `pmid_overrides.yaml` (see `mono_allelic_host` and `ms_samples`).
Rows can legitimately have `is_monoallelic=True` with an empty
`mhc_restriction` — for example, supplementary contaminant peptides under
a mono-allelic PMID override carry the flag but not a per-row allele. If
your downstream pipeline needs a strict "mono-allelic AND has allele"
subset, post-filter on
`is_monoallelic & mhc_restriction.str.startswith("HLA-")`.

### Reports

```bash
hitlist report [--class I|II] [--output report.txt]
```

### QC diagnostics

```bash
hitlist qc                                 # run all checks, print summary
hitlist qc resolution                      # allele-resolution histogram
hitlist qc normalization                   # YAML alleles whose normalize_allele output drifts
hitlist qc mhc-tokens                      # invalid/parser-gap/unknown MHC tokens across modalities
hitlist qc cross-reference                 # alleles in YAML but not data, and reverse
hitlist qc discrepancies [--by sample]     # per-PMID curation drift signals
hitlist qc plan [--top 10]                 # ranked next-PMID-to-curate roadmap
hitlist qc proteome-coverage               # per-source-organism proteome registry coverage
hitlist qc proteome-coverage --missing-only --min-rows 100
```

## Development

```bash
./develop.sh    # install in dev mode
./format.sh     # ruff format
./lint.sh       # ruff check + format check
./test.sh       # pytest with coverage (~3 min)
./deploy.sh     # lint + test + build + upload to PyPI
```

`test.sh` refuses to start pytest when the memory probe fails (exit 2), and
distinguishes this from measured low capacity (exit 1). `--retry-memory` retries
only low capacity. An explicit `TEST_SH_ALLOW_UNKNOWN_MEMORY=1` permits one
worker with an unavailable probe; it does not bypass measured low capacity.

Ordinary tests use fresh data directories and small fixtures, so an installed
corpus cannot inflate their memory use. Full bulk-proteomics and supplementary
data assertions run in the integration phase. The default budgets are 2.5 GiB
per unit worker, and 14 GiB per integration worker on Linux or 24 GiB on macOS.
These include headroom above the Linux 11.6 GiB RSS peak and macOS ~19.5 GB
physical footprint; compression makes macOS RSS alone an underestimate.
Each phase reports its elapsed time
and process peak RSS with `/usr/bin/time` (disable with `TEST_SH_PROFILE=0`).
Peak RSS is not the sum of simultaneous workers' memory. A preflight is a
capacity check; it cannot reserve memory against other programs starting later.

The `CTA` gene set is available in the default install and uses OncoRef's canonical
bundled panel, including its reviewed specificity exclusions. The `cta` extra
remains a compatibility alias. OncoRef is loaded only when querying the gene set.

When local memory cannot support the corpus tests, the **Release build** GitHub
workflow runs the same gates on a hosted runner. It requires the full CI corpus,
installs the current development heads of the six sibling libraries, and records
their resolved revisions with the source commit and artifact checksums. The
workflow builds artifacts; PyPI authentication stays on the maintainer's machine.

From a clean, up-to-date `main` checkout, dispatch the workflow and wait for its
successful manual run. A pull-request validation run cannot authorize publication.

```bash
gh workflow run release-build.yml --ref main
gh run list --workflow release-build.yml
```

Set `release_run_id` to that successful manual run's ID, then download into an
empty directory and verify before uploading. The verifier rejects a different
source commit or version, a non-main checkout, failed or PR workflow runs, and
modified distributions. Publish the downloaded wheel and sdist without rebuilding:

```bash
set -e
release_dir=$(mktemp -d)
gh run download "$release_run_id" \
  --name "hitlist-release-$(git rev-parse HEAD)" --dir "$release_dir"
python scripts/release_artifacts.py verify "$release_dir"
python scripts/check_distribution_license.py "$release_dir"
twine check "$release_dir"/*.whl "$release_dir"/*.tar.gz
twine upload "$release_dir"/*.whl "$release_dir"/*.tar.gz
```

`./deploy.sh --build-only` retains the lint, full test, build and license gates,
and stops before upload. It does not mark a release as published.

See [docs/pmid-curation.md](docs/pmid-curation.md) for the curation YAML format and per-study overrides.
