Metadata-Version: 2.4
Name: lumawarp_py
Version: 0.5.0
Summary: Official Python client and MCP server for Lumawarp, the machine learning platform by Lucidity Sciences.
Author-email: Lucidity Sciences <support@luciditysciences.com>
License-Expression: LicenseRef-Proprietary
Project-URL: Homepage, https://lumawarp.ai
Project-URL: Web app, https://app.lumawarp.ai
Keywords: lumawarp,lucidity-sciences,machine-learning,api-client,mcp,mcp-server
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx>=0.27
Provides-Extra: pandas
Requires-Dist: pandas>=1.5; extra == "pandas"
Provides-Extra: mcp
Requires-Dist: mcp<3,>=2.0; extra == "mcp"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pandas>=1.5; extra == "dev"
Requires-Dist: mcp<3,>=2.0; extra == "dev"
Dynamic: license-file

# lumawarp_py

The official Python client for **Lumawarp**, the machine learning platform by
Lucidity Sciences — the programmatic twin of the web app's day-to-day work.
Sign in, upload and organize datasets, train and manage models, run
inference, fetch and manage predictions, share models, and watch jobs
through to completion. Creating an account and managing your profile stay
in the web app.

- Requires Python 3.10 or newer.
- One runtime dependency: [`httpx`](https://www.python-httpx.org/). pandas
  is optional (`lumawarp_py[pandas]`) and unlocks DataFrame in / DataFrame
  out.
- Knows where the API lives: no base URL to configure.
- Every request is authenticated with the same JWT bearer token the web app
  uses (12-hour expiry).
- Kept in step with the API by a contract test, not by promise — see
  [Staying current](#staying-current).

## Install

From PyPI (import name `lumawarp_py`):

```
pip install lumawarp_py                  # the client
pip install "lumawarp_py[pandas]"        # + DataFrame in / DataFrame out
pip install "lumawarp_py[pandas,mcp]"    # + the MCP server (lumawarp-mcp)
```

From a checkout of this repository: `pip install "./sdk[pandas,mcp]"` (add `-e`
for an editable install). A printable API reference, `LUMAWARP_PY_API.pdf`,
ships in the repository's `docs/` folder and in the distribution kit; regenerate
it after a release with `python sdk/docs/build_api_pdf.py` (needs `reportlab`).
Releases are listed in [CHANGELOG.md](CHANGELOG.md).

## Sign in, sign out

```python
from lumawarp_py import Lumawarp

# Sign in on construction ...
lw = Lumawarp("ada@example.com", "your-password")

# ... or with a token you already hold ...
lw = Lumawarp(token="eyJhbGciOi...")

# ... or construct it signed out and sign in later.
lw = Lumawarp()
lw.login("ada@example.com", "your-password")

lw.me()                 # the signed-in user
lw.is_authenticated     # True
lw.logout()             # discard the token; protected calls now raise NOT_AUTHENTICATED
```

Accounts are created in the web app; the package has no sign-up call.

The client talks to the production API (`lumawarp_py.DEFAULT_BASE_URL`,
`https://api.lumawarp.ai`). To point it at a local stack or staging, pass
`base_url="http://localhost:8000"` or set `LUMAWARP_BASE_URL` for the whole
process. Sessions are stateless JWTs: signing out discards the token
client-side (there is no server-side session to revoke). Tokens expire after
12 hours; call `lw.login(...)` again to refresh.

## Worked example

A runnable version of the core loop — sign in, train from a DataFrame, infer
from a CSV, load the predictions as a DataFrame — lives in
[`examples/quickstart.py`](examples/quickstart.py).

```python
import pandas as pd
from lumawarp_py import Lumawarp, JobStatus, LumawarpError

lw = Lumawarp("ada@example.com", "your-password")

# --- Datasets -----------------------------------------------------------
# Labels in the rightmost column; values may be numeric, text, or missing.
# Upload the path of a .csv file or a pandas DataFrame (nothing else), and
# always say whether the first row is a header row.
manifest = lw.datasets.upload("train.csv", name="housing-2025", has_headers=True)
manifest = lw.datasets.upload(pd.read_csv("train.csv"), name="housing-2026", has_headers=True)
print(manifest["row_count"], manifest["headers"])

lw.datasets.list()                      # newest first
lw.datasets.info("housing-2025")        # full manifest: means, variances, ...
lw.datasets.preview("housing-2025")     # {"headers": [...], "rows": [...]} (first 10)
lw.datasets.rename("housing-2025", "housing-v2")
lw.datasets.delete("housing-v2")

# Folders (organization only - never touches the data)
lw.datasets.create_folder("research")
lw.datasets.set_folder("housing-2025", "research")     # None to unfile
lw.datasets.rename_folder("research", "archive")
lw.datasets.delete_folder("archive")                    # datasets become unfiled

# --- Training, and waiting for it --------------------------------------
job = lw.datasets.train("housing-2025", model_name="housing-model")
print(job["job_id"], job["status_label"])   # e.g. 1765480000123456789 Queued  (job_ts: the same id as an int)

done = lw.jobs.wait(job["job_ts"], timeout=3600, poll_interval=5, raise_on_failure=True)
assert done["status_code"] == JobStatus.COMPLETE

# --- Models -------------------------------------------------------------
lw.models.list()                            # own models + models shared with you
lw.models.info("housing-model")
lw.models.info("m", shared_from="Ada-L")    # inspect a model shared by Ada-L
lw.models.rename("housing-model", "housing-model-v2")
lw.models.delete("housing-model-v2")        # also revokes its shares
# folders: create_folder / rename_folder / delete_folder / set_folder, as for datasets

# --- Inference: CSV or DataFrame in, DataFrame out ----------------------
job = lw.models.infer("housing-model", "batch.csv", has_headers=False)
job = lw.models.infer("housing-model", pd.read_csv("batch.csv", header=None), has_headers=False)
lw.jobs.wait(job["job_ts"])

for p in lw.models.predictions("housing-model"):        # [{filename, size_bytes, last_modified}]
    frame = lw.models.load_prediction("housing-model", p["filename"])         # pandas DataFrame
    lw.models.download_prediction("housing-model", p["filename"], dest=p["filename"])
lw.models.prediction_url("housing-model", "predictions-1.csv")   # presigned URL, 15 minutes
lw.models.delete_prediction("housing-model", "predictions-1.csv")

# --- Sharing ------------------------------------------------------------
lw.models.share("housing-model", recipient="Grace-H")
lw.models.revoke("housing-model", recipient="Grace-H")

# --- Jobs ---------------------------------------------------------------
lw.jobs.list()                              # all, newest first
lw.jobs.list(active=True)                   # Queued / Processing only
lw.jobs.get(job["job_ts"])                  # one job, by id
# each record: {job_ts, job_id, job_type, status_code, status_label, dataset_name, model_name, epoch_ts}

# --- Account & dashboard ------------------------------------------------
lw.account.profile()                        # read here; edit it in the web app
lw.account.billing()                        # plan + pricing, invoices, usage and cost, remaining credit, payment status
lw.account.shares()                         # {"given": [...], "received": [...]}
lw.account.organization()                   # folder document
lw.overview()                               # the Overview page in one call
lw.health()                                 # {"status": "ok", "api_version": "1.0.0"}

lw.close()   # or use the client as a context manager: with Lumawarp(...) as lw:
```

Notes on inputs and return values:

- **Inputs are a `.csv` path or a pandas DataFrame — nothing else.** An open
  file object, another file format, or a path without the `.csv` extension is
  refused before anything is sent (`TypeError` / `ValueError`).
- **`has_headers` is required** on `datasets.upload` and `models.infer`:
  `True` if the first row holds column names (they fill the manifest and are
  never stored with the data), `False` if it is data (the API names the
  columns `Feature_1 ... Target`). There is no default and no sniffing.
- **DataFrames** are sent without their index (so the frame's rightmost
  column is the label) and with NaN as missing cells. `has_headers=True`
  writes the column labels as the header row; `False` drops them — the right
  choice for a frame read with `pd.read_csv(..., header=None)`.
- **A `has_headers` value the data contradicts is accepted, not rejected:** the
  manifest then carries a `warnings` list (for example a text-only first row
  above numeric rows, declared as data and therefore stored as a data row that
  counts in `row_count`), and `datasets.info` returns it again later. Check for
  it after an upload. Dataset and model records also carry `folder` (`None`
  when unfiled, always for shared models).
- Methods return plain dicts and lists decoded from the API's JSON, with the
  documented single-key wrapper (`{"dataset": ...}`, `{"job": ...}`, ...)
  already unwrapped for you.
- **Job records** — from `datasets.train`, `models.infer`, `jobs.list`,
  `jobs.get`, `jobs.wait` and the job lists in `overview()` — carry exactly
  `lumawarp_py.JOB_FIELDS`: `job_ts`, `job_id`, `job_type`, `status_code`,
  `status_label`, `dataset_name`, `model_name` and `epoch_ts` (queued at,
  epoch seconds): what the web app's Overview shows. `job_ts` is the id as an
  integer (exact in Python) and `job_id` the same id as a string: the integer
  is above 2^53, so anything that passes job records through JSON
  (JavaScript, a spreadsheet, an LLM) must keep `job_id`. Queue internals are not
  exposed; the month's usage comes from `account.billing()`, together
  with the contract month's remaining credit (`credit`), the card on file
  (`payment_method`) and `billing_status` — `suspended` means `train` and
  `infer` raise `PAYMENT_REQUIRED` until the overdue invoice is paid.
- `JobStatus` is an `IntEnum` (`QUEUED=0, PROCESSING=1, COMPLETE=2,
  FAILED=3`) that compares equal to the integer `status_code` in job
  records; `.is_terminal` / `.is_active` read as you would expect.
- `jobs.wait` polls `GET /jobs/{job_ts}` every `poll_interval` seconds and
  returns the final job record (`job_ts` or the `job_id` string both work).
  It raises `JOB_TIMEOUT` when `timeout` elapses and, with
  `raise_on_failure=True`, `JOB_FAILED` for a job that ends in failure.
- `jobs.list` returns every job in the live queue, newest first, with no
  paging: the queue is an append-only ledger that holds finished jobs for
  about 24 hours before the coordinator archives them, so the list stays about
  a day deep. Records outlive the datasets and models they name; there is no
  delete.
- `models.rename` copies first and moves the model file last: an error
  (`STORAGE_UNAVAILABLE`) means nothing changed and the old name still
  applies; a `warnings` list on the returned model means the rename succeeded
  and only a follow-up step (shares, old copies) needs another try.
- `models.infer` returns the queued job. If the API attaches advisory
  warnings (for example a column-count mismatch with the training manifest),
  they are emitted through Python's `warnings` module with category
  `lumawarp_py.LumawarpWarning`.
- `load_prediction` (pandas) and `download_prediction` (to disk) both fetch
  a short-lived presigned URL from the API and read the file from it; the
  bearer token is never sent to the object store. `prediction_url` gives
  you the URL itself.
- **Web app only:** creating an account (with its card set-up), editing
  the profile, uploading an avatar and managing the payment method have no
  wrapper by design (`lumawarp_py.WEB_APP_ONLY_ENDPOINTS` lists them, along
  with the Stripe webhook the API exposes for Stripe itself).

## MCP server

The package also runs as an [MCP](https://modelcontextprotocol.io) server, so
an MCP client - Claude Desktop, Claude Code, or an agent built on the Claude
API - can work your Lumawarp account through tools. It is an optional extra;
installing it changes nothing about the client.

Nothing to install with [uv](https://docs.astral.sh/uv/) or pipx:

```
uvx --from "lumawarp_py[mcp]" lumawarp-mcp
pipx run --spec "lumawarp_py[mcp]" lumawarp-mcp
```

or install it and run the `lumawarp-mcp` command (`lumawarp-mcp --version`
checks the install without starting the server):

```
pip install "lumawarp_py[mcp]"
claude mcp add lumawarp -e LUMAWARP_EMAIL=you@example.com -e LUMAWARP_PASSWORD=... -- lumawarp-mcp
claude mcp add lumawarp -e LUMAWARP_EMAIL=you@example.com -e LUMAWARP_PASSWORD=... -- uvx --from "lumawarp_py[mcp]" lumawarp-mcp
```

Or in a client's JSON configuration (Claude Desktop and most others):

```json
{ "mcpServers": { "lumawarp": {
    "command": "uvx", "args": ["--from", "lumawarp_py[mcp]", "lumawarp-mcp"],
    "env": { "LUMAWARP_EMAIL": "you@example.com", "LUMAWARP_PASSWORD": "..." } } } }
```

- Runs locally over stdio and signs in once from the environment; credentials
  never pass through tool arguments. `LUMAWARP_EMAIL` + `LUMAWARP_PASSWORD` is
  the convenient form (an expired token is renewed automatically);
  `LUMAWARP_TOKEN` (a 12-hour session token, from `Lumawarp(...).token`) keeps
  the password out of the config file for a bounded session.
- Tools mirror the client: `list_datasets`, `upload_dataset`, `train_model`,
  `run_inference`, `list_predictions`, `get_prediction`, `wait_for_job`,
  `get_billing`, ... (`lumawarp_py.mcp_server.TOOL_SOURCES` lists every tool
  with the client method behind it). Deletes and revokes carry MCP's
  destructive annotation so clients can ask before running them.
- Uploads take a local `.csv` path or the CSV text (written to a temporary
  `.csv`, so the same rules apply) and always require `has_headers`.
- **Job ids are strings over MCP.** Every job record carries `job_id` (and
  `job_ts` rendered the same way) as a string of digits, because the ids exceed
  the range a JSON number carries exactly; `get_job` and `wait_for_job` take
  that string back unchanged and refuse a rounded number.
- **`wait_for_job` waits.** It blocks until the job is Complete or Failed,
  reporting progress on every poll, and one call is capped at an hour; pass
  `max_wait_seconds` to bound it yourself. An unfinished answer carries
  `next_step` telling the model to call again with the same `job_id` rather
  than stop, and `train_model` / `run_inference` return `next_step` too.
- Jobs are an append-only ledger: no delete tool, records outlive their
  datasets and models, and finished jobs leave `list_jobs` after about a day
  when they are archived.
- Read-only resources: `lumawarp://datasets`, `lumawarp://models`,
  `lumawarp://jobs`, `lumawarp://jobs/{job_id}` and `lumawarp://billing`.
- No sign-up, profile editing or avatar tools - the same web-app-only rule.
- The SDK suite checks that every tool rests on an existing client method,
  that every client method is reachable or deliberately excluded, and that
  importing `lumawarp_py` never loads `mcp`, so the server can neither drift
  from the package nor weigh it down.

## Error handling

Every failed call raises `lumawarp_py.LumawarpError`, parsed from the API's
standard error envelope:

```json
{ "detail": { "code": "DATASET_EXISTS", "message": "You already have a dataset named 'housing-2025'." } }
```

```python
from lumawarp_py import Lumawarp, LumawarpError

try:
    lw.datasets.upload("train.csv", name="housing-2025", has_headers=True)
except LumawarpError as err:
    print(err.code)      # "DATASET_EXISTS"
    print(err.message)   # "You already have a dataset named 'housing-2025'."
    print(err.status)    # 409
```

Codes you will encounter include `UNAUTHORIZED`, `INVALID_CREDENTIALS`,
`CSV_REJECTED`, `DATASET_NAME_INVALID`, `DATASET_EXISTS`,
`DATASET_NOT_FOUND`, `MODEL_NAME_INVALID`, `MODEL_EXISTS`, `MODEL_NOT_FOUND`,
`PREDICTION_NOT_FOUND`, `JOB_NOT_FOUND`, `RECIPIENT_NOT_FOUND`,
`CANNOT_SHARE_WITH_SELF`, `ALREADY_SHARED`, `SHARE_NOT_FOUND`,
`ACTIVE_JOB_CONFLICT`, `FOLDER_NOT_FOUND`, `STORAGE_UNAVAILABLE` and `INTERNAL`.
`INVALID_CREDENTIALS` reads the same for an unknown email and a wrong password
(account existence is never revealed); its message points at *Forgot
password?* on the sign-in page at app.lumawarp.ai.

Client-side codes cover situations the server never saw:

| Code | Meaning |
|---|---|
| `NOT_AUTHENTICATED` | A protected method was called on a signed-out client; nothing was sent. |
| `JOB_TIMEOUT` | `jobs.wait` gave up before the job finished. |
| `JOB_FAILED` | `jobs.wait(..., raise_on_failure=True)` saw the job end in failure. |
| `VALIDATION_ERROR` | The server returned FastAPI's default list-shaped 422 detail; the messages are joined into `err.message`. |
| `HTTP_ERROR` | The response was not the standard envelope (plain-string detail, HTML from a proxy, empty body). |
| `NETWORK_ERROR` | The request never completed (connection failure, timeout); `err.status` is `None`. |

A `LumawarpError` prints as `CODE: message (HTTP status)`, so bare
`except LumawarpError as err: print(err)` is already informative.

Mistakes in how a method is called are caught before any request and raise
ordinary Python exceptions: `TypeError` (data that is neither a `.csv` path
nor a DataFrame; a missing or non-boolean `has_headers`) and `ValueError` (a
path that does not end in `.csv`; a non-positive `poll_interval`).

## Staying current

This package tracks the API mechanically:

- Every method that wraps an API operation is registered with the operation
  it covers (`lumawarp_py.wrapped_endpoints()` lists them).
- `backend/tests/test_sdk_contract.py` — part of the backend suite that
  gates every deploy — diffs that registry against the live app's OpenAPI
  schema, so an endpoint added without a wrapper (or a wrapper left behind
  after an endpoint is removed) fails the build. The only gap it allows is
  `lumawarp_py.WEB_APP_ONLY_ENDPOINTS` (sign-up and its card set-up, the
  Stripe webhook, profile editing, avatar upload, the billing portal), and
  that list must stay exact. The same file drives this client
  through the real app across the whole lifecycle, DataFrames included.
- `lumawarp_py.API_VERSION` names the API contract this release was built
  against; the API reports its own as `api_version` on `GET /health`, and
  the contract test pins the two together.

## Using the API without Python (curl)

The SDK is a thin veneer over the HTTP API (base path `/api/v1`). The same
workflow with curl:

```sh
BASE="https://api.lumawarp.ai/api/v1"

# Log in and capture the token
TOKEN=$(curl -s -X POST "$BASE/auth/login" \
  -H "Content-Type: application/json" \
  -d '{"email": "ada@example.com", "password": "your-password"}' \
  | python -c "import json,sys; print(json.load(sys.stdin)['token'])")

AUTH="Authorization: Bearer $TOKEN"

# Upload a dataset (multipart: file + form fields "name" and "has_headers")
curl -s -X POST "$BASE/datasets" -H "$AUTH" \
  -F "file=@train.csv;type=text/csv" -F "name=housing-2025" -F "has_headers=true"

# List datasets / inspect / preview
curl -s "$BASE/datasets" -H "$AUTH"
curl -s "$BASE/datasets/housing-2025" -H "$AUTH"
curl -s "$BASE/datasets/housing-2025/preview" -H "$AUTH"

# Queue a training job, then poll it
curl -s -X POST "$BASE/datasets/housing-2025/train" -H "$AUTH" \
  -H "Content-Type: application/json" -d '{"model_name": "housing-model"}'
curl -s "$BASE/jobs/1765480000123456789" -H "$AUTH"

# Models and inference
curl -s "$BASE/models" -H "$AUTH"
curl -s "$BASE/models/housing-model" -H "$AUTH"
curl -s -X POST "$BASE/models/housing-model/infer" -H "$AUTH" \
  -F "file=@batch.csv;type=text/csv" -F "has_headers=false"

# Predictions: list, then fetch the presigned URL and download from it
curl -s "$BASE/models/housing-model/predictions" -H "$AUTH"
URL=$(curl -s "$BASE/models/housing-model/predictions/out.csv" -H "$AUTH" \
  | python -c "import json,sys; print(json.load(sys.stdin)['download_url'])")
curl -s -o out.csv "$URL"   # presigned: no Authorization header here
curl -s -X DELETE "$BASE/models/housing-model/predictions/out.csv" -H "$AUTH"

# Share and revoke
curl -s -X POST "$BASE/models/housing-model/share" -H "$AUTH" \
  -H "Content-Type: application/json" -d '{"recipient_username": "Grace-H"}'
curl -s -X DELETE "$BASE/models/housing-model/share/Grace-H" -H "$AUTH"

# Jobs, profile, billing, health
curl -s "$BASE/jobs?active=true" -H "$AUTH"
curl -s "$BASE/account/profile" -H "$AUTH"
curl -s "$BASE/account/billing" -H "$AUTH"
curl -s "$BASE/health"
```

## Development

```
pip install -e "sdk[dev]"
cd sdk && pytest                                   # unit suite, mock transport
cd backend && pytest tests/test_sdk_contract.py    # live contract against the real app
cd sdk && python -m build && python -m twine check --strict dist/*   # release artifacts
```

Releases go to PyPI from a `sdk-v<version>` tag; the runbook is in
[`docs/CICD.md`](../docs/CICD.md) ("Publishing the Python package").
