Metadata-Version: 2.4
Name: gtfreader
Version: 0.3.0
Summary: Fast Cython-backed parsing for GTF attribute columns.
Author: Endre Bakken Stovner
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Cython
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Operating System :: OS Independent
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.12
Description-Content-Type: text/markdown
Requires-Dist: pandas>=2.0
Requires-Dist: numpy
Provides-Extra: fast-io
Requires-Dist: pyarrow; extra == "fast-io"

# gtfreader

Fast GTF reading into pandas DataFrames.

## Install

```bash
python -m pip install -e .
```

### Optional: parse on every core

pandas' CSV parser is single-threaded, and on a large GTF the nine fixed
columns are about a third of the cost of reading. With `pyarrow` installed that
part is parsed on every core instead:

```bash
python -m pip install -e ".[fast-io]"
```

`read_gtf` is 1.5x faster end to end at 10^6 rows on twelve cores. It is only
the tabular parse that speeds up; expanding the attribute column is the rest of
the work and is unchanged.

pandas remains the reference implementation and the fallback. The pyarrow path
is skipped when pyarrow is missing, when `nrows` is set, when a file is shaped
in a way it does not model, and whenever the two parsers would disagree -- the
result is the same frame either way.

## Example

```python
from gtfreader import read_gtf

df = read_gtf("annotation.gtf")
print(df.columns)
print(df.head())
```
