Metadata-Version: 2.4
Name: streamxl
Version: 5.3.2
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Financial and Insurance Industry
Classifier: Natural Language :: English
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Rust
Classifier: Topic :: Office/Business
Classifier: Topic :: Utilities
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Dist: rich>=13.0
Requires-Dist: pytest>=7.0 ; extra == 'dev'
Requires-Dist: pytest-cov>=4.0 ; extra == 'dev'
Requires-Dist: flask>=2.0 ; extra == 'dev'
Requires-Dist: openpyxl>=3.0 ; extra == 'dev'
Requires-Dist: ruff>=0.2.0 ; extra == 'dev'
Requires-Dist: black>=24.1.0 ; extra == 'dev'
Requires-Dist: flask>=2.0 ; extra == 'server'
Provides-Extra: dev
Provides-Extra: server
License-File: LICENSE
Summary: High-performance .xlsx file reader with formula extraction, comment support, and cross-sheet dependency analysis. Streams large .xlsx files row-by-row without loading the entire file into memory; 3.8-5.6x faster than openpyxl in this repo's own committed benchmarks (see benchmarks/results.md). Rust-powered, Python API.
Keywords: excel,xlsx,spreadsheet,streaming,file-parsing,data-loading,performance,rust,python-api,memory-efficient,big-data,etl,data-pipeline,file-handling,openpyxl-alternative,pandas-io,data-engineering,batch-processing,file-processing,performance-optimization
Author-email: Georgi Mammen Mullassery <mullassery@gmail.com>
Maintainer-email: Georgi Mammen Mullassery <mullassery@gmail.com>
License-Expression: Apache-2.0
Requires-Python: >=3.8
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Bug Tracker, https://github.com/Mullassery/PyStreamXL/issues
Project-URL: Changelog, https://github.com/Mullassery/PyStreamXL/releases
Project-URL: Discussions, https://github.com/Mullassery/PyStreamXL/discussions
Project-URL: Documentation, https://github.com/Mullassery/PyStreamXL#readme
Project-URL: Homepage, https://github.com/Mullassery/PyStreamXL
Project-URL: Repository, https://github.com/Mullassery/PyStreamXL
Project-URL: Source Code, https://github.com/Mullassery/PyStreamXL/tree/main

# StreamXL

## Problem

`openpyxl` and similar pure-Python Excel readers load the whole workbook
into memory before you can touch a single row — fine for small files, a
real ceiling for ETL pipelines and data engineering workloads working
against large `.xlsx` exports.

## Solution

**Stream large `.xlsx` files row-by-row in constant memory, powered by a Rust core with no `unsafe` code.**

`pip install`s as `streamxl`, `import streamxl`. Read multi-sheet Excel workbooks without loading them fully into memory, extract formulas and comments, write new `.xlsx` files, and append to existing ones — all through a small, plain Python API backed by a Rust engine.

[![PyPI](https://img.shields.io/pypi/v/streamxl)](https://pypi.org/project/streamxl/)
[![CI](https://github.com/Mullassery/PyStreamXL/actions/workflows/ci.yml/badge.svg)](https://github.com/Mullassery/PyStreamXL/actions/workflows/ci.yml)
[![Python 3.8+](https://img.shields.io/badge/Python-3.8%2B-blue)](https://www.python.org)
[![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](./LICENSE)

## Use cases

- **ETL against large Excel exports** that don't fit comfortably in memory
  with `openpyxl` — `read()` keeps memory flat regardless of file size.
- **Extracting formulas/comments for audit or migration tooling**, not
  just cell values.
- **Appending to a growing log-style `.xlsx` file** without rewriting the
  whole workbook or losing other sheets.
- **Not yet a good fit for:** SQL-style querying across sheets, formula
  *evaluation* (only extraction/classification), or pandas/Parquet/Arrow
  export built in — see [Honest feature list](#honest-feature-list) for
  the full "what's not here" list.

---

## Install

```bash
pip install streamxl
```

A prebuilt wheel is currently published only for macOS (arm64); other platforms install from the source distribution, which requires a Rust toolchain (see `rust-toolchain.toml`) and [maturin](https://www.maturin.rs/) to build. Every PyPI release to date (1.2.0 through 5.2.0) has shipped exactly one platform wheel plus an sdist — no Linux or Windows wheels have been published yet.

## Quick start

```python
import streamxl

for row in streamxl.read("data.xlsx"):
    print(row)  # ['Name', 'Age', 'Score']
```

`read()` streams rows one at a time — memory use stays flat regardless of file size.

## Real, working examples

**Read as dictionaries, keyed by header row:**

```python
import streamxl

for row in streamxl.read("sales.xlsx", as_dict=True):
    print(row["Customer"], row["Amount"])
```

**Read only specific columns:**

```python
for row in streamxl.read("sales.xlsx", as_dict=True, columns=["Customer", "Amount"]):
    ...
```

**Read every sheet in a workbook:**

```python
sheet_names = streamxl.sheets("workbook.xlsx")
all_data = streamxl.read_all("workbook.xlsx")  # {sheet_name: [rows...]}
```

**Write a new `.xlsx` file:**

```python
import datetime
import streamxl

streamxl.write("report.xlsx", [
    ["Name", "Joined", "Score"],
    ["Alice", datetime.date(2024, 1, 15), 95.5],
    ["Bob", datetime.date(2024, 3, 2), 88.0],
])
```

**Stream-write multiple sheets without holding the whole file in memory:**

```python
with streamxl.writer("report.xlsx") as w:
    w.write_row(["Name", "Age"])
    w.write_row(["Alice", 30])
    w.add_sheet("Summary")
    w.write_row(["Total", 1])
```

**Append rows to an existing file (other sheets are preserved):**

```python
streamxl.write("log.xlsx", [["Date", "Event"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "started"]])
streamxl.append("log.xlsx", [[datetime.date.today(), "finished"]])
```

**Extract formulas and comments:**

```python
rows = list(streamxl.read("model.xlsx", with_formulas=True))
# each cell is a dict: {"value": ..., "formula": ..., "formula_type": ...,
#                        "comment": ..., "comment_author": ...}

from streamxl import FormulaSerializer
export = FormulaSerializer.export_formulas(rows)
FormulaSerializer.export_to_json(rows, "formulas.json")
FormulaSerializer.export_to_csv(rows, "formulas.csv")  # sanitized against CSV/formula injection
```

**Export to CSV safely** — untrusted cell content is never written to CSV verbatim (see [Security](#security) below):

```python
import csv
import streamxl
from streamxl.security import sanitize_csv_cell

with open("output.csv", "w", newline="") as f:
    writer = csv.writer(f)
    for row in streamxl.read("large.xlsx"):
        writer.writerow([sanitize_csv_cell(cell) for cell in row])
```

**Validate a file and recover from bad cells instead of crashing:**

```python
from streamxl import validate_excel_file

report = validate_excel_file("questionable.xlsx")
if report.has_fatal_errors():
    print(report.format_summary())
```

More runnable examples live in [`examples/`](examples/).

## Honest feature list

What's here and real, backed by the Rust core and covered by the test suite:

- **Streaming reads** — `read()` / `stream()`: a real Rust `__iter__`/`__next__` iterator over the sheet — you get rows one at a time, not a pre-built Python list. **Correction (2026-09-22 benchmark):** this is *not* O(1) memory as previously claimed here. `XlsxStream::open()` decompresses the whole sheet XML into memory before iteration starts, so peak RSS scales with sheet size (measured ~1.3MB of RSS per 1MB of `.xlsx`, real data — see [`benchmarks/results.md`](benchmarks/results.md)). It's still far below `openpyxl`'s full-load mode and the API shape (a lazy iterator) is real, but memory is not flat regardless of file size — tracked as gap #10 in [`ROADMAP_HONEST.md`](ROADMAP_HONEST.md). `read_rows_all_at_once()`/`read_rows_with_metadata_all_at_once()` remain available as an explicit escape hatch for callers that need random access or to iterate the result more than once.
- **Multi-sheet support** — `sheets()`, `read_all()`, and `writer().add_sheet()`.
- **Streaming writes** — `write()`, `writer()`, `append()`, all producing real `.xlsx` files.
- **Formula extraction** — read formula text and a best-effort formula-type classification (`with_formulas=True`), plus `FormulaReferenceMapper` for shifting/rewriting cell references and `FormulaSerializer` for exporting/importing formulas as JSON or CSV.
- **Comment extraction** — cell comments and authors, via `with_formulas=True`.
- **Conditional formatting rules** — `conditional_formats()` reads every `<conditionalFormatting>`/`<cfRule>` in a sheet (type, operator, formulas, priority, `stopIfTrue`) and resolves each rule's `dxfId` against `xl/styles.xml`'s `<dxfs>` into concrete font color/bold/italic and fill colors. `colorScale`/`dataBar`/`iconSet` rules are captured (type, sqref, priority) but their inline color-stop/threshold definitions aren't modeled — those rule types don't use `dxfId` in the first place.
- **Type-aware cells** — strings, numbers, booleans, dates, datetimes, and empty cells round-trip correctly.
- **Error recovery & validation** — `validate_excel_file()` and `ErrorRecoveryHandler` classify and (optionally) recover from malformed cells instead of hard-failing on the whole file.
- **Security hardening** — path validation, file-size limits, and ZIP-bomb defenses (entry-size, compression-ratio, and total-decompressed-size limits) enforced before/while a file is opened. CSV export is sanitized against formula-injection (see below).
- **REST API (optional)** — `streamxl.server.StreamXLServer` / `create_flask_app()` wrap the real streaming engine behind HTTP endpoints (`/sources`, `/sources/<id>/query`, `/sources/<id>/export`, ...). Requires `pip install "streamxl[server]"`.

What's **not** here, so you don't have to find out the hard way:

- No SQL-style query language — `execute_query()` in the REST API streams rows from a named sheet, it does not parse arbitrary queries.
- No pandas/Parquet/Arrow export built in. Convert `read()`'s output yourself, or open an issue if this matters to you.
- No formula *evaluation* — formula text is extracted and classified, not recalculated.
- The `pystreamxl dashboard` CLI command renders sample data, not live telemetry — every mode (bare, `--static`, `--alerts`, `--recommendations`, `--export`) shows the same explicit "SAMPLE DATA — not live" warning.

## Security

- **Path & size validation** — `validate_read_path()` / `validate_write_path()` reject non-`.xlsx` paths, path traversal, and oversized files before any parsing happens.
- **ZIP-bomb defenses** — the Rust core enforces a per-entry size limit, a compression-ratio limit, and a total-decompressed-size limit while unpacking a workbook (see `core/src/zip_reader.rs`), tested against real crafted archives in `core/tests/zip_bomb_defense.rs`.
- **CSV/formula-injection protection** — `streamxl.security.sanitize_csv_cell()` neutralizes any string cell that starts with `=`, `+`, `-`, `@`, TAB, or CR (the standard CSV-injection trigger set) by prefixing it with `'`, so a malicious workbook can't turn a CSV export into an executable formula when reopened in Excel/LibreOffice/Google Sheets. `FormulaSerializer.export_to_csv()` applies this automatically; apply it yourself when writing CSV from `read()` output (see the example above).

**Limits, enforced by default (no configuration needed):**

| Limit | Value |
|---|---|
| Max file size | 512 MB |
| Max size per ZIP entry | 512 MB |
| Max total decompressed size | 1 GB |
| Max compression ratio | 30:1 |

Handle malformed or malicious files by catching `SecurityError`:

```python
from streamxl import SecurityError, read

try:
    for row in read("data.xlsx"):
        process(row)
except SecurityError as e:
    print(f"Security violation: {e}")
```

Found a security issue? See [SECURITY.md](SECURITY.md).

## Performance

Rows are parsed and yielded one at a time rather than being collected into a Python list up front, and it's consistently faster than `openpyxl` on both reads and writes. Memory use is lower than `openpyxl`'s full-load mode but currently scales with sheet size rather than staying flat — see the correction under "Honest feature list" above and [`ROADMAP_HONEST.md`](ROADMAP_HONEST.md) gap #10. See [`benchmarks/`](benchmarks/) for the scripts used to compare against `openpyxl`, and [`examples/memory_benchmark.py`](examples/memory_benchmark.py) to measure it yourself against your own files:

```bash
python examples/memory_benchmark.py your_file.xlsx
```

Actual numbers depend heavily on your file's structure (shared strings, formulas, formatting) — measure on your own workloads rather than trusting a generic table.

### vs openpyxl, on real data

Methodology: 150,000 real, live NYC 311 Service Request rows pulled from
NYC Open Data's Socrata API (`data.cityofnewyork.us/resource/erm2-nwe9`,
current as of 2026-09-22 — not synthetic/fabricated rows), 14 columns,
written to a real 2-sheet `.xlsx` workbook (75k rows/sheet, 19MB) via
`openpyxl`. Both libraries iterated every row of every sheet; row counts
and a positional checksum matched exactly across all three methods
(correctness verified, not just speed). 3 runs each, median reported,
single-process wall-clock via `time.perf_counter()`, peak RSS via
`resource.getrusage(...).ru_maxrss` on macOS/arm64, Python 3.13.

| Rows | streamxl `read()` | openpyxl `read_only=True` | openpyxl full load |
|------|---|---|---|
| 10,000  | 0.07s · 28MB peak RSS | 0.61s · 31MB peak RSS | — |
| 30,000  | 0.19s · 52MB peak RSS | 1.89s · 32MB peak RSS | — |
| 75,000  | 0.48s · 103MB peak RSS | 4.72s · 36MB peak RSS | — |
| 150,000 (2 sheets) | 0.96s · 192MB peak RSS | 9.27s · 43MB peak RSS | 13.5s · 1,040MB peak RSS |

**streamxl is ~9.7x faster than `openpyxl(read_only=True)` and ~14x
faster than `openpyxl()` full-load** at 150k rows — but at that size it
uses **~4.5x more peak memory than `openpyxl(read_only=True)`** (192MB
vs 43MB), because `read()` isn't actually O(1) yet (see above). If your
bottleneck is wall-clock time, streamxl wins clearly. If your bottleneck
is memory on a very large file and you don't need every column loaded at
once, `openpyxl(read_only=True)` currently uses less RAM. Reproduce with
`benchmarks/openpyxl_vs_streamxl.py` against any real `.xlsx` file.

## CLI

```bash
pystreamxl dashboard          # sample extraction dashboard (unlabeled placeholder, see note above)
pystreamxl dashboard --static # same sample data, clearly labeled "SAMPLE DATA — not live"
pystreamxl --version
```

## Development

```bash
git clone https://github.com/Mullassery/PyStreamXL.git
cd PyStreamXL
pip install -e ".[dev]"       # builds the Rust extension via maturin and installs test deps
pytest tests/ -v
cargo test --release --all-features     # Rust unit + integration tests (both core and python crates)
```

On macOS you may need `RUSTFLAGS="-C link-args=-undefined -C link-args=dynamic_lookup"` before `cargo build`/`cargo test` for the PyO3 extension crate to link outside of `maturin`/`pip install`.

See [CONTRIBUTING.md](CONTRIBUTING.md) before opening a PR.

## Docs

- [`docs/architecture/README.md`](docs/architecture/README.md) — how the Rust engine and Python API fit together, including known dead code
- [`docs/xlsx_format.md`](docs/xlsx_format.md) — XLSX/ZIP/XML format notes
- [`ROADMAP_HONEST.md`](ROADMAP_HONEST.md) — unvarnished list of what's missing, broken, or technical debt
- [`CHANGELOG.md`](CHANGELOG.md) — release history
- [`SECURITY.md`](SECURITY.md) — security model, limits, and what it does *not* protect against

## License

This project is licensed under the [Apache License 2.0](LICENSE).

---

**StreamXL** | Constant-memory Excel streaming | Rust core, Python API

