Metadata-Version: 2.3
Name: gwenlake
Version: 0.9.6
Summary: Gwenlake Python Client
Requires-Dist: click>=8.1.7
Requires-Dist: pyarrow>=14.0
Requires-Dist: httpx>=0.28.1
Requires-Dist: pandas>=3.0.3
Requires-Dist: pydantic>=2.12.5
Requires-Dist: python-dotenv>=1.2.1
Requires-Python: >=3.11
Description-Content-Type: text/markdown

# Gwenlake Python Library

The Gwenlake Python library provides convenient access to the Gwenlake API
from applications written in Python. A single `Gwenlake` client gives you access
to your catalog — projects, datasets, files and SQL.

## Installation

```sh
pip install -U gwenlake
```

Or install the latest development version straight from GitHub:

```sh
pip install -U git+https://github.com/gwenlake/gwenlake-python
```

## Authentication

The client authenticates with a Bearer token, resolved in this order:

1. an explicit `api_key` / `credentials` passed to the client,
2. a named `profile`,
3. the `GWENLAKE_API_KEY` environment variable,
4. the `default` profile in `~/.gwenlake/credentials`.

```bash
export GWENLAKE_API_KEY='sk-...'
```

```python
from gwenlake import Gwenlake

# uses GWENLAKE_API_KEY, or the default ~/.gwenlake/credentials profile
client = Gwenlake()

# or pass the key explicitly
client = Gwenlake(api_key="sk-...")

# or pick a profile from ~/.gwenlake/credentials
client = Gwenlake(profile="myteam")
```

The `~/.gwenlake/credentials` file is an INI file with one section per profile,
holding either a static `token` (API key) or OAuth2 `client_id` / `client_secret`.

## Projects

```python
projects = client.projects.list()
for p in projects:
    print(p["alias"], p["id"])

project = client.projects.get("res.project.…")
```

## Datasets

```python
datasets = client.datasets.list()
for d in datasets:
    print(d["alias"], d["id"])

dataset = client.datasets.get("res.dataset.…")
```

## Files

Files live inside a dataset.

```python
dataset_id = "res.dataset.…"

# list files
for f in client.files.list(dataset_id):
    print(f["filename"], f["file_size"])

# upload a local file (optionally into a subdirectory with path=...)
client.files.upload(dataset_id, "report.pdf")
client.files.upload(dataset_id, "report.pdf", path="docs")

# download a file
content = client.files.download(dataset_id, "report.pdf")

# presigned URL / delete
url = client.files.presigned_url(dataset_id, "report.pdf")
client.files.delete(dataset_id, "report.pdf")
```

## SQL

Run SQL against a dataset (DuckDB), referencing it as
`'<project_alias>.<dataset_alias>'`. With `format="json"` the rows are returned
under `data`:

```python
result = client.statements.create(
    statement="SELECT * FROM 'flights.flight-data' LIMIT 10",
    format="json",
)
for row in result["data"]:
    print(row)
```

Pass a `connection_id` to run the statement against a connection's native engine
(PostgreSQL, S3, …) instead of a dataset.

## Transforms

A Palantir Foundry-style transforms layer (`gwenlake.transforms`) lets you write
dataset-to-dataset transformations as decorated functions. Datasets (and
models) are addressed as `"<project_alias>.<alias>"` — the same handle used in
SQL.

`transform_df` — the function receives each `Input` as a `pandas.DataFrame` and
**returns** the DataFrame to write to the (single) `Output`. The result is
written automatically (snapshot/replace by default):

```python
from gwenlake.transforms import transform_df, Input, Output

@transform_df(
    raw_data=Input("Project_A.users"),
    processed_data=Output("Project_A.users_filtered"),
)
def process(raw_data):
    df = raw_data[raw_data["age"] >= 18].copy()
    df["name_upper"] = df["name"].str.upper()
    return df

process(client)   # reads, computes, writes
```

`transform` — the lower-level form: the function receives `TransformInput` /
`TransformOutput` objects and reads/writes explicitly. Use it for non-tabular
data (images, PDFs, …) via `.filesystem()`:

```python
from gwenlake.transforms import transform, Input, Output

@transform(
    my_input=Input("Project_A.users"),
    my_output=Output("Project_A.users_distinct"),
)
def dedupe_users(my_input, my_output):
    df = my_input.dataframe()
    # mode="replace" (default) clears the dataset first; "append" keeps existing files
    my_output.write_dataframe(df.drop_duplicates(), mode="replace")

@transform(
    images=Input("Project_A.scans"),
    thumbnails=Output("Project_A.scans_processed"),
)
def process_files(images, thumbnails):
    src, dst = images.filesystem(), thumbnails.filesystem()
    for entry in src.ls():
        data = src.read(entry["filename"])          # raw bytes (PDF, image, …)
        with dst.open(f"copy/{entry['filename']}", "wb") as f:
            f.write(data)
```

### Models

A **model** is a catalog resource whose artifacts live in the git repository the
code lives in (`/models` in the catalog). `train` produces one, and `Model(...)`
binds one a transform loads — conventionally as `model=`:

```python
import joblib
from gwenlake.transforms import train, transform_df, Input, Model, Output

@train(
    training_set=Input("Project_A.churn_training"),
    output=Output("Project_A.churn"),          # an Output of @train is a MODEL
)
def fit(training_set, output):
    clf = fit_classifier(training_set)         # training_set is a DataFrame
    joblib.dump(clf, output.file("model.pkl")) # write under the model's directory
    return {"auc": 0.91}                       # returned dict -> the model's metrics

@transform_df(
    customers=Input("Project_A.customers"),
    model=Model("Project_A.churn"),            # a model this transform loads
    output=Output("Project_A.churn_scores"),
)
def predict(customers, model):
    clf = joblib.load(model.file("model.pkl"))
    return customers.assign(churn=clf.predict(customers))
```

Models are a catalog resource served by api-catalog. Its `/models` endpoints
are **not** routed through the public gateway today (that path serves the
inference model list), so `model.info()` / `model.update()` work from inside a
build — where the client already points at api-catalog — but not against
`api.gwenlake.com`. `model.path` never needs an API call during a build.

`output.path` (and `model.path`) is the model's directory **in the checkout**:
during a build the engine sets it, commits whatever the training run wrote
there, and pins that commit as the model's version. `model.parameters`,
`model.version` and `model.update(metrics=..., version=...)` cover the model
card. Lineage follows: `datasets -> train -> model -> transform -> dataset`.

**Large datasets** — page through with `LIMIT/OFFSET` instead of loading
everything at once. `iter_dataframes()` yields `pandas.DataFrame` chunks and
`write_dataframes()` streams them back out as `part-00000.parquet`, …:

```python
@transform(
    big_dataset=Input("Project_A.events"),
    result=Output("Project_A.events_clean"),
)
def transform_in_chunks(big_dataset, result):
    chunks = (
        chunk[chunk["valid"]]
        for chunk in big_dataset.iter_dataframes(chunk_size=50_000, order_by="id")
    )
    result.write_dataframes(chunks, mode="replace")
```

Pass `order_by=` for a deterministic page split. The transforms layer is
synchronous.

## Async

Every resource is also available on `AsyncGwenlake`:

```python
import asyncio
from gwenlake import AsyncGwenlake

async def main():
    client = AsyncGwenlake()
    print(await client.projects.list())

asyncio.run(main())
```

See [`examples/`](examples/) for runnable scripts.
