Metadata-Version: 2.4
Name: twobitreader-rs
Version: 0.2.2
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Development Status :: 4 - Beta
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Programming Language :: Python :: Free Threading :: 2 - Beta
Classifier: Typing :: Typed
Requires-Dist: pytest ; extra == 'tests'
Provides-Extra: tests
License-File: LICENSE-MIT
License-File: LICENSE-APACHE
Summary: Python bindings for the twobitreader Rust crate
License-Expression: MIT OR Apache-2.0
Requires-Python: >=3.9
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Repository, https://github.com/andrewdelong/twobitreader-rust

# twobitreader-rs

Fast DNA sequence extraction from [2bit files](http://genome.ucsc.edu/FAQ/FAQformat.html#format7),
the compact genome format used by the UCSC Genome Browser.

The motivation for this package is speed. See the underlying [`twobitreader`](https://crates.io/crates/twobitreader) Rust crate for benchmarks.

```console
pip install twobitreader-rs
```

## Usage

```python
from twobitreader_rs import TwobitReader

tbr = TwobitReader.open("hg38.2bit")
tbr.get("chr1", 10000, 10005)        # 'TAACC'
tbr.seq_len("chr1")                  # 248956422
tbr.names()                          # ['chr1', 'chr2', ...]
```

Ranges are 0-based with an exclusive end, like a Python `start:end` slice.
Every method also has an `_inclusive` counterpart taking 1-based inclusive ranges,
as used by GFF/GTF and genome browsers.

**Concatenation** helps to assemble a spliced transcript from its exons:

```python
from twobitreader_rs import reverse_complement

tx = tbr.concat("chr6", [(1389575, 1391118), (1394695, 1395603)])
tx = reverse_complement(tx)          # if the transcript is on the minus strand
```

**Prefetching** dramatically speeds up access to "cold" files that are not yet in
the page cache:

```python
tbr.prefetch(exons)                  # tell the OS what is coming
seqs = tbr.get_batch(exons)          # reads land as the data arrives
```


## Documentation

The Python methods closely mirror their counterparts in the [Rust crate's documentation](https://docs.rs/twobitreader).
Every Python method carries a full docstring with arguments, return values and exceptions, visible
through `help(TwobitReader)`.

## License

MIT OR Apache-2.0

