Metadata-Version: 2.4
Name: invoicedataextraction-sdk
Version: 0.6.1
Summary: Official Python SDK for Invoice Data Extraction.
Author-email: Invoice Data Extraction <developers@invoicedataextraction.com>
License-Expression: MIT
Project-URL: Homepage, https://invoicedataextraction.com
Project-URL: Documentation, https://invoicedataextraction.com/sdk/python
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: requests>=2.28.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Dynamic: license-file

# invoicedataextraction-sdk

Official Python SDK for [Invoice Data Extraction](https://invoicedataextraction.com). Uploads your files, submits the extraction, waits for it to finish and hands you the rows as data, or a spreadsheet, in a few lines of code.

- Python 3.9 or later

## Install

```bash
pip install invoicedataextraction-sdk
```

## Quick Start

```python
import json
import os
import sys

from invoicedataextraction import InvoiceDataExtraction
from invoicedataextraction.errors import SdkError, ApiResponseError

try:
    client = InvoiceDataExtraction(
        api_key=os.environ.get("INVOICE_DATA_EXTRACTION_API_KEY"),
    )

    result = client.extract(
        folder_path="./invoices",
        prompt="Extract invoice number, date, vendor name, and total amount",
        output_structure="per_invoice",
        json_typed_values=True,
        console_output=True,  # remove to disable console logging
    )

    if result["status"] == "completed":
        for row in client.iterate_results(extraction_id=result["extraction_id"]):
            print(row)  # one dict per extracted row, keyed by your output columns
except (SdkError, ApiResponseError) as error:
    print(json.dumps(error.body, indent=2), file=sys.stderr)
    raise SystemExit(1)
```

`extract(...)` uploads your files (pass a `folder_path` or a list of `files`), submits the extraction, waits until it finishes and returns the result: the final status response from the API, for a completed, failed or cancelled extraction. `iterate_results(...)` then reads the extracted rows straight from the API, with amounts as numbers and empty cells as `None` because the extraction was submitted with `json_typed_values`; `get_results(...)` reads one page when you want to manage paging yourself. Check `result["pages"]["failed_count"]` to verify that all uploaded pages were processed, and `result["review_needed"]["count"]` for rows that need a human's check before you rely on the data.

To get a spreadsheet, add `download={"formats": ["xlsx"], "output_path": "./output"}` to the call and the file is saved when the extraction completes, or call `download_output(...)` later.

Generate an API key from your [dashboard](https://invoicedataextraction.com/dashboard?view=API). Every account includes 50 free pages per month. Additional credits can be purchased on a pay-as-you-go basis with no subscription needed.

## Staged Workflow

If you need control over individual steps, for example uploading files in one part of your system and extracting in another, use the lower-level methods:

```python
import json
import os
import sys

from invoicedataextraction import InvoiceDataExtraction
from invoicedataextraction.errors import SdkError, ApiResponseError

try:
    client = InvoiceDataExtraction(
        api_key=os.environ.get("INVOICE_DATA_EXTRACTION_API_KEY"),
    )

    upload = client.upload_files(
        files=["./invoice1.pdf", "./invoice2.pdf"],
        console_output=True,
    )

    submitted = client.submit_extraction(
        upload_session_id=upload["upload_session_id"],
        file_ids=upload["file_ids"],
        prompt="Extract invoice number and total",
        output_structure="per_invoice",
        json_typed_values=True,
    )

    result = client.wait_for_extraction_to_finish(
        extraction_id=submitted["extraction_id"],
        console_output=True,
    )

    for row in client.iterate_results(extraction_id=submitted["extraction_id"]):
        print(row)
except (SdkError, ApiResponseError) as error:
    print(json.dumps(error.body, indent=2), file=sys.stderr)
    raise SystemExit(1)
```

## Options, and stopping a run

Beside `json_typed_values`, `extract(...)` and `submit_extraction(...)` take `output_language`, `review_needed_fill_color`, `affected_field_fill_color` and `send_completion_email`, each applying to that extraction only. Without `json_typed_values`, every value in the JSON output and in the rows is a string.

Set `ask_questions=True` and the extraction can stop to ask when the documents leave something unsettled, instead of deciding on its own. Pass `on_questions` and the SDK calls it as `on_questions(questions, status)`, sends back the answers it returns and carries on to the result; without it, `extract(...)` returns `status: "input_required"` with the questions for you to answer with `answer_questions(...)`. The questions, the answer forms and the deadline are in the [API reference](https://invoicedataextraction.com/api#input-required).

```python
def answer(questions, status):
    return [
        {"question_id": question["question_id"], "accept_recommended": True}
        # or {"question_id": ..., "choice_id": "b"}, or {"question_id": ..., "text": "DD/MM/YYYY"}
        for question in questions
    ]

result = client.extract(
    folder_path="./invoices",
    prompt="Extract invoice number, date, vendor name, and total amount",
    output_structure="per_invoice",
    json_typed_values=True,
    ask_questions=True,
    on_questions=answer,
)

# Stop an extraction that is still queued or processing
client.cancel_extraction(extraction_id=extraction_id)
```

## Listing past extractions

```python
# One-page browse with filters
page = client.list_extractions(
    status="completed",
    limit=50,
)

# Auto-paginating iterator over every matching extraction
for extraction in client.iterate_extractions(status="completed"):
    print(extraction["extraction_id"], extraction["task_name"])

# Full record (the original prompt, options, full pages, full failure error, etc.)
result = client.get_extraction(extraction_id="...")
extraction = result["extraction"]
```

On listing methods, team admins can pass `scope="team"` to see extractions submitted by any team member; pass `scope="own"` to force own-only results.

## Error Handling

SDK methods raise `SdkError` or `ApiResponseError` on failure. The structured error body is on `error.body`, with fields `error.body["error"]["code"]`, `error.body["error"]["message"]`, `error.body["error"]["retryable"]`, and `error.body["error"]["details"]`.

When an extraction task itself reaches a terminal state, `extract(...)` returns that response rather than raising: check `result["status"]` for `"completed"`, `"failed"`, or `"cancelled"` for tasks stopped from the web app or with `cancel_extraction(...)`. See the [full docs](https://invoicedataextraction.com/sdk/python) for details.

## Documentation

- [Python SDK docs](https://invoicedataextraction.com/sdk/python): full method reference, parameters, return shapes, and examples
- [REST API docs](https://invoicedataextraction.com/api): endpoint-level documentation for direct HTTP integration
- [Dashboard](https://invoicedataextraction.com/dashboard?view=API): manage API keys and view extraction results

## License

MIT
