Metadata-Version: 2.1
Name: pggt
Version: 1.4.0
Summary: Graph-to-allele, CNV and PAV workflow for multi-genome analysis
Author: PGGT developers
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Natural Language :: English
Classifier: Natural Language :: Chinese (Simplified)
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: psutil >=5.9

# PGGT 1.4.0

`PGGT` is a recoverable multi-genome workflow that converts pangenome graph
projections into allele, copy-number variation (CNV), and presence/absence
variation (PAV) components.

It runs six modules:

1. project annotated genes onto GFA walks and build initial components;
2. filter graph-supported components with DIAMOND and gene structure evidence;
3. score component genes using CDS, exon, Pfam and optional expression evidence;
4. regroup singleton genes with reciprocal Miniprot and Pfam evidence;
5. write member tables, an integer PAV/CNV matrix and two W-only GFA views;
6. write both all-gene and protein-coding GFF3 models for each sample.

Module 3 uses component-specific score thresholds: `1.0` for `singleton`,
`0.9` for `one_to_one_partial`, and `0.8` for `one_to_one_allele` and
`one_to_many_supported`. A gene is retained as `PASS` regardless of its score
when its CDS is complete, Pfam status is `PASS`, and total exon length is
strictly greater than 200 bp. Other genes below their component threshold
remain in the main component table with `ng` replacing the terminal gene
marker `g`. The original-to-output relationship and decision are recorded in
`13.gene_id_mapping.tsv`.

## What's new in 1.4.0

- Expression input is optional. Omit `-e / --expression-dir` to skip only
  expression scoring; other scores, thresholds and strong-evidence rescue remain.
- Each sample gets `<sample>.all_genes.gff3` and
  `<sample>.protein_coding.gff3`. Noncoding IDs retain lowercase `ng`.
- `PAV_CNV.copy_number.tsv` combines presence/absence and actual copy number:
  `0` absent, `1` one gene, `2, 3, 4, ...` the number of gene copies.
- W-only GFA is generated separately for all genes and PASS genes; repeated
  component visits retain distinct copies.
- Exports validate coverage, gene hierarchy, matrix counts and W-path counts.
  Source GFF genes must all occur in the classification mapping and final
  components; unresolved differences stop export with a coverage report.

`NOCODING` is a computational classification, not an experimentally validated
RNA biotype. Original biotypes do not automatically override classification.
Source CDS and other features remain in all-gene exports for NOCODING genes.

## Installation

PGGT requires Linux and Python 3.9 or later.

```bash
pip install pggt
```

The workflow also calls external command-line programs that are not installed
by pip:

- `bedtools`
- `diamond`
- `samtools`
- HMMER (`hmmscan` and `hmmpress`)
- `miniprot`

For example, these programs can be prepared with Conda/Bioconda:

```bash
conda create -n pggt -c conda-forge -c bioconda \
  python=3.10 bedtools diamond samtools hmmer miniprot
conda activate pggt
pip install pggt
```

Confirm that the command is available:

```bash
pggt --help
```

## Quick start

Run input validation first. `--dry-run` creates no result files.

```bash
pggt \
  --gfa-input pangenome.gfa \
  --genome-dir genomes/ \
  --gff-dir annotations/ \
  --samples species_list.txt \
  --pfam-hmm Pfam-A.hmm \
  --output-dir pggt_out/ \
  --threads 56 \
  --parallel-jobs 4 \
  --dry-run
```

This example omits expression input. Add `--expression-dir expression/` (or
`-e expression/`) when expression tables are available. After validation
succeeds, run the same command without `--dry-run`.

To reuse Pfam work when rerunning the same species and annotations, add either
an earlier PGGT result directory or its compact Pfam table:

```bash
pggt [all required arguments] \
  --previous-pfam-results previous_pggt_out/
```

The previous and current runs must use a Pfam database with the same
fingerprint. Reuse is sequence-based, so proteins absent from the previous
table or with changed SHA256 values are scanned automatically.

PGGT 1.1 tables do not retain the statistics for every accepted hit. To avoid
carrying the 1.1 best-hit direction bug forward, unambiguous no-domain and
single-domain legacy rows are reusable; legacy multi-domain rows use an exact
shared-cache record when available and otherwise are rescanned. Tables created
by PGGT 1.2 or later retain the corrected best hit and are fully reusable.

The output directory must be new. Do not use a GraphTribe result directory as
the PGGT output directory.

## Input requirements

### Sample list

`species_list.txt` must contain exactly one sample name per non-empty line
and at least two unique samples:

```text
B73
EIL02
EIL31
EIL08
EIL37
```

Sample names may contain letters, numbers, underscores, dots, and hyphens. The
first sample is the reference sample used for representative-gene selection.
Every listed sample must occur in the second field of at least one GFA `W`
record.

Sample names resolve genome and GFF files, and expression files when enabled.
Each enabled input type must resolve uniquely for every sample.

### GFA

`--gfa-input` accepts either one GFA file or a directory containing exactly
one `*.gfa` file. Use `--gfa-name FILE` when a directory contains multiple
GFA files.

GFA `W` coordinates are interpreted as 0-based, half-open intervals.

### Genome FASTA files

Provide one uncompressed genome FASTA per sample. Recommended names are:

```text
genomes/B73.fa
genomes/EIL02.fa
```

PGGT first checks `<sample>.fa`, `<sample>.fasta`, `<sample>.fna`, and
`<sample>.fas`. If none exists, it accepts one uniquely matching FASTA whose
name contains the sample name.

FASTA sequence names must match the chromosome or sequence names used by the
corresponding GFF and GFA walks.

### GFF/GFF3 files

Recommended names are:

```text
annotations/B73.renamed.sorted.chr.gff3
annotations/EIL02.renamed.sorted.chr.gff3
```

PGGT also checks `<sample>.gff3` and `<sample>.gff`, followed by one unique
matching GFF file whose name contains the sample name.

The annotations must provide gene, transcript, exon, and CDS relationships.
Gene IDs must match expression tables when supplied and use the supported
terminal `g` plus digits pattern, such as `B73_C05g06442`; this allows PGGT to
mark a noncoding output ID as `B73_C05ng06442` without ambiguity. Chromosome
names must match the corresponding genome FASTA.

### Expression tables

Expression input is optional for the run. If supplied, the directory must contain
one file for every sample; invalid paths, missing files and invalid tables remain
errors. Each file must be a
tab-separated TSV table. The recommended file name is
`<sample>.FPKM.tsv`; `<sample>.expression.tsv` is also accepted.

Example:

```text
Gene_ID	leaf_FPKM	root_FPKM
B73_C01g00001	12.50	3.20
B73_C01g00002	0	0.08
B73_C01g00003	0	0
```

Expression-table rules:

- The header must contain the exact, case-sensitive column name `Gene_ID`.
- At least one expression column must end exactly with `_FPKM`.
- Gene IDs must match gene IDs in the corresponding GFF/GFF3.
- FPKM cells must be numeric. Do not write `NA`, `N/A`, or other text.
- Empty FPKM cells are interpreted as zero.
- When multiple `*_FPKM` columns are present, PGGT uses the maximum value for
  each gene.
- Use one row per gene. If a gene ID occurs more than once, the later row
  replaces the earlier value.
- Genes absent from a supplied table receive zero expression points with
  `expression_status=GENE_NOT_FOUND`.

Extra non-FPKM columns are allowed. PGGT first checks
`<sample>.FPKM.tsv` and `<sample>.expression.tsv`; otherwise, exactly one
`*<sample>*.tsv` file must exist in the expression directory. Ambiguous
matches stop the run.

Without expression input, the total is `CDS_score + exon_score + Pfam_score`.
With expression input, PGGT adds the existing `expression_score`. There is no
renormalization, extra penalty or threshold change when expression is omitted.
Other evidence can still produce PASS, including strong coding-evidence rescue.

The retained `score_tmp_file/05.expression_scores.tsv` contains audit rows even
when scoring is skipped. Total-score rows also contain `expression_status`:

| Situation | max_FPKM | expression_score | expression_status |
| --- | --- | --- | --- |
| No expression directory | NA | 0.00 | SKIPPED_NO_INPUT |
| Gene missing from supplied table | NA | 0.00 | GENE_NOT_FOUND |
| Measured zero | 0.000000 | 0.00 | OBSERVED |
| Positive measurement | Observed value | Existing rule | OBSERVED |

The module-3 shell entry accepts optional `-e / --expression`; its Python entry
accepts optional `--expression-dir`. Partial per-sample expression availability
is not supported in this release.

### Pfam database

`--pfam-hmm` must point to a non-empty `Pfam-A.hmm` file. PGGT prepares the
required HMMER index files in its working area when necessary.

## Run control

```bash
# Resume a previously interrupted PGGT run
pggt [all original arguments] --resume

# Run only one module
pggt [all required arguments] --only-step 3

# Run a module range
pggt [all required arguments] --from-step 2 --to-step 4
```

Use the same inputs, output directory, thresholds, thread count, and job count
when resuming. Completed modules are validated through their completion
markers before being skipped. Modules 5 and 6 additionally check the new
exports and their counts, including valid empty coding files.

Use a new output directory when upgrading from 1.3.0, changing expression mode,
or changing expression inputs. Resume checks the version, expression enablement,
directory and content fingerprint and rejects mismatches. Compatible Pfam
evidence can still be reused through `--previous-pfam-results`.

## Threads and shared caches

`--threads` is the total CPU budget. `--parallel-jobs` controls concurrent
per-sample work, Pfam shards, Miniprot mappings, and PAF parsing, and must not
exceed `--threads`.

For five samples on a sufficiently large 56-core node, `--parallel-jobs 5`
allows all five Miniprot target mappings to run concurrently. Values above the
sample count can still increase Pfam shard concurrency, but should be
benchmarked against available memory and storage bandwidth.

Shared caches are stored under `~/.cache/pggt`:

- `~/.cache/pggt/pfam/pfam_cache.sqlite3` caches Pfam results by HMM
  fingerprint and protein sequence SHA256;
- `~/.cache/pggt/miniprot/` stores Miniprot genome indexes.

Set the `PGGT_CACHE_DIR` environment variable to use a different cache root.

## Main outputs

```text
pggt_out/
├── 00.run_info/                         # parameters, logs, timings
├── 00.input_view/                       # expression links only if enabled
├── status/
├── 01.graph_mapping/
├── 02.component_filter/
│   └── 10.filtered_component_table.txt
├── 03.gene_score/
│   ├── 11.scored_gene_component_table.txt
│   ├── 12.nocoding_gene_scores.tsv
│   ├── 13.gene_id_mapping.tsv
│   └── score_tmp_file/
│       ├── 05.expression_scores.tsv
│       ├── 07.pfam_scores.tsv
│       ├── 08.singleton_score.tsv
│       └── 09.summary.txt
├── 04.component_cnv/
│   └── 04.orthogroup/
│       └── singleton_orthogroup_component_table.txt
├── 05.allele_CNV_PAV/
│   ├── allele_CNV_PAV_component_table.txt
│   ├── PAV_CNV.copy_number.tsv
│   ├── protein_coding_component_table.txt
│   ├── representative_gene_selection.tsv
│   ├── protein_coding_representative_selection.tsv
│   ├── gene_coverage.tsv
│   ├── module5_metadata.tsv
│   ├── artifact_manifest.tsv
│   └── W_gfa/
│       ├── <sample>.W.gfa               # compatibility copy: all genes
│       ├── allele_CNV_PAV.W.gfa         # compatibility copy: all genes
│       ├── all_genes/
│       │   ├── <sample>.all_genes.W.gfa
│       │   ├── all_genes.W.gfa
│       │   └── <view audit tables>
│       └── protein_coding/
│           ├── <sample>.protein_coding.W.gfa
│           ├── protein_coding.W.gfa
│           └── <view audit tables>
└── 06.coding_gene_gff/
    ├── <sample>.protein_coding.gff3
    ├── <sample>.all_genes.gff3
    ├── <sample>.gene_coverage.tsv
    ├── coding_gene_gff_summary.tsv
    ├── all_gene_gff_summary.tsv
    └── artifact_manifest.tsv
```

Each GFA view contains `component_segment_map.tsv`, `gene_segment_map.tsv`,
`chromosome_w_path_stats.tsv`, `missing_or_extra_genes.tsv` and
`make_w_gfa.summary.tsv`. No PASS genes produces an empty coding W file and
status `NO_CODING_GENES`.

The integer matrix counts unique gene members including `ng`, not transcripts;
sample columns follow the sample list. Member tables retain IDs for GFA export.
Coding components keep the original component ID and class label while their
membership, gene count and sample count are filtered. If a representative is
noncoding, the coding view reselects a reproducible PASS representative. Join
views by component ID; node names can differ. These are W-only gene-order
paths, not complete sequence graphs with S/L records.

Each completed pipeline module contains `artifact_manifest.tsv`. Large intermediate
DIAMOND, Pfam, and Miniprot files are not published into successful result
directories. Failed module workspaces remain under `OUTPUT_DIR/.work/` for
diagnosis.

## Standalone module 5

When a compatible module-4 component table already exists, run module 5
directly:

```bash
pggt-module5 \
  --component-table singleton_orthogroup_component_table.txt \
  --samples-file species_list.txt \
  --gff-dir annotations/ \
  --gene-id-map 13.gene_id_mapping.tsv \
  --output-dir module5_out/ \
  --representative-seed 1
```

`--gene-id-map` is now always required, including for all-PASS input, because
both views require explicit classification. Required columns are `sample`,
`original_gene_id`, `output_gene_id` and `decision`. Duplicate memberships or
missing classifications are errors. `gene_coverage.tsv` reports unresolved
GFF/mapping/component differences.

## Standalone module 6

Module 6 can be run directly from a module-3 mapping table:

```bash
pggt-module6 \
  --gene-id-map 13.gene_id_mapping.tsv \
  --samples-file species_list.txt \
  --gff-dir annotations/ \
  --output-dir coding_gene_gff/ \
  --parallel-jobs 4
```

Both GFF views are written in one run. Protein-coding output selects
`decision=PASS` and preserves original IDs. All-gene output includes PASS and
NOCODING, uses `output_gene_id` (lowercase `ng` for NOCODING), and adds
`PGGT_decision` and `PGGT_original_gene_id` attributes to gene rows.

Both views retain original transcript, exon, CDS and UTR models, coordinates,
strand and phase. Child IDs remain unchanged; Parent and geneID references are
updated for renamed or filtered genes. Multi-parent features and children before
parents are supported; invalid hierarchies and rename collisions are rejected.
Source row order is retained. Embedded FASTA is omitted.

The supported gene root feature is `gene`. Other gene-like roots such as
`pseudogene` require upstream normalization and are rejected with a diagnostic.
The original GFF and mapping must cover the same gene set. Per-sample coverage
reports identify missing entries; this release does not invent classifications
or singleton components for uncovered genes.

## Command overview

```bash
pggt --help
pggt-module5 --help
pggt-module6 --help
```

## Maintainer notes

Before a PyPI release, run the regression tests, build both the wheel and
source distribution, run `twine check dist/*`, and verify installation in a
clean environment. Update this README whenever command-line options or input
requirements change.

Major workflow changes must be developed in a new versioned package directory,
with matching updates to `pyproject.toml`, runtime `__version__`, this
README, `PGGT流程原理.md`, and `CHANGELOG.md`. Do not modify a previously
published version directory in place.

## License

License information will be added before the first public release.
