Metadata-Version: 2.4
Name: jasentool
Version: 1.3.0
Summary: Multipurpose tool for jobs related to the jasen pipeline and Bonsai tool.
Author-email: Ryan Kennedy <Ryan.Kennedy@skane.se>
Project-URL: Repository, https://github.com/ryanjameskennedy/jasentool
Project-URL: Issues, https://github.com/ryanjameskennedy/jasentool/issues
Project-URL: Changelog, https://github.com/ryanjameskennedy/jasentool/CHANGELOG.md
Project-URL: Documentation, https://jasentool.readthedocs.io
Classifier: Development Status :: 3 - Alpha
Classifier: Natural Language :: English
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: setuptools
Requires-Dist: wheel
Requires-Dist: requests
Requires-Dist: tqdm
Requires-Dist: pandas
Requires-Dist: pymongo<4,>=3.12
Requires-Dist: openpyxl
Requires-Dist: biopython
Requires-Dist: matplotlib
Requires-Dist: seaborn
Requires-Dist: click>=8.1
Requires-Dist: pysam>=0.22
Requires-Dist: cyvcf2
Requires-Dist: numpy>=1.24
Requires-Dist: rauth
Requires-Dist: pyyaml
Requires-Dist: pydantic<3,>=2.4
Provides-Extra: dev
Requires-Dist: pylint~=4.0; extra == "dev"
Requires-Dist: black~=26.3; extra == "dev"
Requires-Dist: isort~=8.0; extra == "dev"
Provides-Extra: test
Requires-Dist: pytest>=7.0; extra == "test"
Requires-Dist: pytest-cov>=4.1; extra == "test"
Dynamic: license-file

# Jasentool

Multipurpose tool for jobs related to the [JASEN](https://github.com/Clinical-Genomics-Lund/JASEN) pipeline and [Bonsai](https://github.com/Clinical-Genomics-Lund/bonsai).

Full documentation: [jasentool.readthedocs.io](https://jasentool.readthedocs.io).

## Installation

```
pip install jasentool
```

### Older Linux distributions (recommended: conda)

On hosts with **glibc < 2.28** (Ubuntu < 18.04, RHEL/CentOS < 8, Debian < 10), pip can't find binary wheels for current pandas and numpy on Python 3.12, so it falls back to source builds that need GCC 9.3 or newer. Use conda instead; conda-forge ships compatible binaries:

```
git clone https://github.com/SMD-Bioinformatics-Lund/jasentool.git
cd jasentool
conda env create -f environment.yml
conda activate jasentool
```

`environment.yml` pulls the heavy dependencies (pandas, numpy, matplotlib, biopython, pysam, cyvcf2, openpyxl) from conda-forge, then installs jasentool itself in editable mode.

`compare-distances` uses [`cgmlst-dists`](https://github.com/tseemann/cgmlst-dists) (from bioconda) to build its distance matrices, and `environment.yml` includes it. It isn't a pip dependency, so `pip install jasentool` won't bring it in; get it through conda or build it yourself. Without it, `compare-distances` falls back to a slower pure-Python calculation.

## Usage

```
jasentool <subcommand> [options]
```

Run `jasentool --help` to list subcommands, or `jasentool <subcommand> --help` for per-subcommand help.

### Subcommands

**Post-run analysis**

| Subcommand | Description |
|------------|-------------|
| `check-backup` | Cross-check Bonsai samples against the backup storage tree |
| `rebuild-manifests` | Rebuild Bonsai manifests from the backup storage tree |
| `rerun-chewbbaca` | Re-run chewBBACA AlleleCall on a check-backup masked-assemblies CSV |
| `compare-distances` | Build cgMLST distance matrices for two chewBBACA tables and their difference |
| `find` | Query samples from MongoDB |
| `identify-missing` | Identify samples absent from JASEN results directory |
| `validate-pipelines` | Compare pipeline outputs against MongoDB records |

**Pipeline processes**

| Subcommand | Description |
|------------|-------------|
| `annotate-delly` | Annotate Delly structural-variant VCFs with gene symbols and locus tags |
| `concatenate-files` | Concatenate multiple YAML files (e.g. `versions.yml`) |
| `count-reads` | Count reads in FASTQ file(s) |
| `create-blacklist` | Aggregate minority base frequencies across BAMs to produce a blacklist TSV |
| `create-yaml` | Create YAML input file for Bonsai upload |
| `format-cdm` | Build a CDM input file from a sample manifest |
| `minority-report` | Compute minority base frequency distribution from a `samtools mpileup` file |
| `post-align-qc` | Compute post-alignment QC from BAM |

**Site-specific hooks**

| Subcommand | Description |
|------------|-------------|
| `reformat-csv` | Reformat BJORN CSV/SH files for JASEN |

**Setup & reference data**

| Subcommand | Description |
|------------|-------------|
| `converge-catalogues` | Merge WHO, TBdb, and FoHM TB mutation catalogues |
| `download-bigsdb` | Download cgMLST scheme alleles from PubMLST or BIGSdb |
| `download-ncbi` | Download genome FASTA and GFF from NCBI |
| `transform-file-format` | Convert cgMLST target TSV to BED format |

## Quick examples

### Query samples from MongoDB

```
jasentool find \
  --query MySampleID \
  --db-name mydb \
  --db-collection samples \
  --output-file results.json
```

### Identify missing samples

```
jasentool identify-missing \
  --output-file missing.json \
  --db-name mydb \
  --db-collection samples \
  --analysis-dir /path/to/jasen/results
```

### Validate pipeline outputs

```
jasentool validate-pipelines \
  --input-dir /path/to/new/results \
  --output-dir /path/to/validation/output \
  --db-name mydb \
  --db-collection samples
```

### Cross-check backup storage

```
jasentool check-backup \
  --profile staphylococcus_aureus \
  --backup-dir /backup/jasen \
  --db-name bonsai \
  --db-collection samples \
  --address mongodb://bonsai.host:27017/ \
  -o backup_status.csv
```

### Compute post-alignment QC

```
jasentool post-align-qc \
  --sample-id SAMPLE_ID \
  --bam-file SAMPLE.bam \
  --output-file SAMPLE_qc.json \
  [--bed-file regions.bed] \
  [--cpus 4]
```

See the [Usage docs](https://jasentool.readthedocs.io/en/latest/usage.html) for full details.
