Metadata-Version: 2.5
Name: narwhals-datafusion
Version: 0.1.3
Summary: Apache DataFusion backend for Narwhals, via the Narwhals plugin system
Project-URL: Repository, https://github.com/s5dsn-eqee/narwhals-datafusion
License-Expression: MIT
License-File: LICENSE.md
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Python: >=3.10
Requires-Dist: datafusion<55,>=54
Requires-Dist: narwhals>=2.25
Requires-Dist: pyarrow>=14.0.0
Provides-Extra: extra-functions
Requires-Dist: datafusion-extra-functions-ffi>=0.1; extra == 'extra-functions'
Description-Content-Type: text/markdown

# narwhals-datafusion

[![CI](https://github.com/s5dsn-eqee/narwhals-datafusion/actions/workflows/ci.yml/badge.svg)](https://github.com/s5dsn-eqee/narwhals-datafusion/actions/workflows/ci.yml)
[![PyPI version](https://badge.fury.io/py/narwhals-datafusion.svg)](https://badge.fury.io/py/narwhals-datafusion)
[![PEP 740](https://img.shields.io/badge/PEP%20740-attested-3775A9?logo=pypi&logoColor=white)](https://pypi.org/project/narwhals-datafusion/#files)
[![Downloads](https://static.pepy.tech/badge/narwhals-datafusion/month)](https://pepy.tech/project/narwhals-datafusion)

[Apache DataFusion](https://datafusion.apache.org/python/) backend for
[Narwhals](https://github.com/narwhals-dev/narwhals), registered through the
`narwhals.plugins` entry point.

```sh
pip install narwhals-datafusion                    # core
pip install "narwhals-datafusion[extra-functions]" # + mode, skew, kurtosis
```

## Usage

```python
import narwhals as nw
import pyarrow as pa
from datafusion import SessionContext

ctx = SessionContext()
df = ctx.from_arrow(pa.table({"a": [1, 2, 3], "b": ["x", "y", "x"]}))

lf = nw.from_native(df)  # nw.LazyFrame on this backend
result = (
    lf.group_by("b")
    .agg(nw.col("a").sum())
    .sort("b")
    .collect(backend="pyarrow")
)
```

Everything is lazy until `.collect()`: expressions become `datafusion.Expr`,
frame verbs become `datafusion.DataFrame` methods, and DataFusion executes the
plan. Dtypes are pyarrow end-to-end.

## Architecture

A subclass of narwhals' SQL layer (`narwhals._sql`, `narwhals._compliant`),
the layer DuckDB, Ibis and Spark also use. Tested against the narwhals release
vendored as the `narwhals/` submodule (2.25) and datafusion 54, the major
`pyproject.toml` pins.

| Module | Class |
|---|---|
| `dataframe.py` | `DataFusionLazyFrame`: frame verbs over `datafusion.DataFrame` |
| `expr.py` | `DataFusionExpr`: the `SQLExpr` hooks, aggregates, windows, casts |
| `namespace.py` | `DataFusionNamespace`: `SQLNamespace` primitives, IO, horizontal functions |
| `group_by.py`, `selectors.py`, `expr_str/dt/list/struct.py` | supporting surface |
| `utils.py` | function-name remapping, window/sort builders, dtype bridge via `narwhals._arrow` |

## `mode`, `skew`, `kurtosis`

These aggregates live in the Rust-only
[`datafusion-extra-functions`](https://github.com/datafusion-contrib/datafusion-extra-functions)
crate. The `extra-functions` extra installs
[`datafusion-extra-functions-ffi`](https://github.com/s5dsn-eqee/datafusion-extra-functions-ffi),
a prebuilt wheel exposing them through datafusion-python's
`__datafusion_aggregate_udf__` capsule protocol. Without it the three methods
raise `NotImplementedError`. The wheel pins the datafusion major it was built
for.

## API coverage

narwhals 2.25 on datafusion 54. ⚠️ entries work with the caveat in
parentheses; see [Known limitations](#known-limitations-as-of-datafusion-54).

| Namespace | ✅ Supported | ⚠️ Partial | ❌ Not supported |
|---|---|---|---|
| `Expr` | `abs` `alias` `all` `any` `any_value` `ceil` `clip` `cos` `count` `cum_count` `cum_max` `cum_min` `cum_sum` `diff` `exp` `fill_nan` `first` `floor` `is_between` `is_close` `is_duplicated` `is_finite` `is_first_distinct` `is_in` `is_last_distinct` `is_nan` `is_null` `is_unique` `last` `len` `log` `max` `mean` `median` `min` `null_count` `over` `pipe` `rank` `rolling_mean` `rolling_std` `rolling_sum` `rolling_var` `round` `shift` `sin` `sqrt` `std` `sum` `var` | `cast` (no `Enum`) · `fill_null` (no `strategy` + `limit`) · `kurtosis` (`[extra-functions]` extra) · `mode` (`[extra-functions]` extra, `keep="any"` only) · `n_unique` (not over windows) · `replace_strict` (explicit `default` required) · `skew` (`[extra-functions]` extra) | `cum_prod` `quantile` |
| `Expr.str` | `contains` `ends_with` `head` `len_chars` `pad_end` `pad_start` `replace_all` `slice` `split` `starts_with` `strip_chars` `strip_chars_end` `strip_chars_start` `tail` `to_lowercase` `to_time` `to_uppercase` `zfill` | `to_date`/`to_datetime` (explicit `format` required) · `to_titlecase` (no word breaks on digits) | `replace` (use `replace_all`) |
| `Expr.dt` | `convert_time_zone` `date` `day` `hour` `microsecond` `millisecond` `minute` `month` `nanosecond` `ordinal_day` `second` `to_string` `truncate` `weekday` `year` | `replace_time_zone` (`None`/`"UTC"` only) | `offset_by` `timestamp` `total_microseconds` `total_milliseconds` `total_minutes` `total_nanoseconds` `total_seconds` |
| `Expr.list` | `contains` `get` `len` `max` `min` `sort` | `unique` (`maintain_order=False` only) | `mean` `median` `sum` |
| `Expr.struct` | `field` | | |
| `LazyFrame` | `collect` `collect_schema` `drop` `drop_nulls` `filter` `group_by` `head` `join` `rename` `select` `sort` `top_k` `unique` `unpivot` `with_columns` `with_row_index` | `explode` (single column) · `sink_parquet` (file path only) | `join_asof` |

Not listed: methods narwhals does not support on any lazy backend
(`Expr.filter`, `Expr.drop_nulls`, `Expr.unique`, `Expr.map_batches`,
`Expr.ewm_mean`, `LazyFrame.tail`, `LazyFrame.gather_every`).

## Known limitations (as of datafusion 54)

- `join_asof`, exact `quantile`, `cum_prod`, `list.sum/mean/median`,
  `dt.total_*`, `dt.offset_by`, `dt.timestamp`, `str.replace`, `Enum` casts:
  no engine support, raise `NotImplementedError`.
- `mode`, `skew`, `kurtosis`: need the `extra-functions` extra.
- `n_unique().over(...)` raises: DataFusion ignores `DISTINCT` inside window
  aggregates, which would return wrong results. It also drops `ORDER BY` and
  `IGNORE NULLS` there; this backend moves those onto the window itself.
- `fill_null(strategy=..., limit=n)` raises: bounded frames with
  `first_value`/`last_value` need `retract_batch`, not implemented engine-side.
- `replace_time_zone` supports `None` and `"UTC"` only; use `convert_time_zone`
  for instant-preserving conversions.
- `str.to_datetime`/`to_date` require an explicit `format`.
- `replace_strict` requires an explicit `default`.
- `nw.scan_csv`/`nw.scan_parquet` cannot dispatch to a plugin backend yet
  (narwhals gap); read with a `SessionContext` and pass the frame to
  `nw.from_native`.
- `str.to_titlecase` uses `initcap`, which does not break words on digits.
- Row order is guaranteed only after `sort`: `concat` may interleave inputs and
  backward `fill_null` may reorder rows.

## Development

```sh
git submodule update --init                       # narwhals at the tested tag
uv sync --group tests --extra extra-functions
uv run --group tests pytest tests                 # this package's tests
uv run --group tests python run_tests.py          # narwhals' suite, known failures deselected
```

See [CONTRIBUTING.md](CONTRIBUTING.md).
