Metadata-Version: 2.5
Name: langchain-flatmark
Version: 1.0.1
Summary: Document to Markdown API and MCP server for PDF, Word, PowerPoint, Excel and HTML. OCR queue for large files. Hosted in Germany.
Project-URL: Homepage, https://flatmark.dev
Project-URL: Documentation, https://github.com/flatmark-dev/flatmark-integrations/blob/main/langchain/README.md
Project-URL: Repository, https://github.com/flatmark-dev/flatmark-integrations
Project-URL: Issues, https://github.com/flatmark-dev/flatmark-integrations/issues
Author-email: podshalocef <contact@podshalocef.com>
License-Expression: MIT
Keywords: document-loader,flatmark,langchain,markdown
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: httpx<1,>=0.28
Requires-Dist: langchain-core<2,>=1.6
Description-Content-Type: text/markdown

# langchain-flatmark

Document to Markdown API and MCP server for PDF, Word, PowerPoint, Excel and HTML. OCR queue for large files. Hosted in Germany.

A LangChain document loader for the [flatmark API](https://flatmark.dev): `FlatmarkLoader` turns local files and URLs into Markdown `Document`s.

```sh
pip install langchain-flatmark
```

```python
from langchain_flatmark import FlatmarkLoader

docs = FlatmarkLoader("report.pdf").load()
print(docs[0].page_content)
```

Direct conversion (`POST /v1/convert`) works without a key at a lower rate limit. [Get an API key](https://flatmark.dev/go/langchain?to=/app/api-keys) and pass it as `api_key=` or set `FLATMARK_API_KEY`.

For large or scanned files, convert through the queue (an API key is required):

```python
loader = FlatmarkLoader(["scan.pdf", "https://example.com/deck.pptx"], use_queue=True)
for doc in loader.lazy_load():
    print(doc.metadata, len(doc.page_content))
```

The loader submits `POST /v1/convert/jobs`, polls `GET /v1/jobs/{job_id}` until the job's status is `succeeded` or `failed`, then downloads `GET /v1/convert/jobs/{job_id}/result`.

## Documents

One `Document` per source; `page_content` is the Markdown. `metadata` carries `source` (the path or URL as given) plus the `meta` fields of the answer — and `job_id` for a queued conversion.

A URL is downloaded by the loader without your key, then uploaded. Other options: `base_url`, `poll_interval`, `timeout`, and `client` (an `httpx.Client` for proxies or retries).

API reference: https://flatmark.dev/docs · Support: https://flatmark.dev/support · Generated from [`openapi.json`](../openapi.json).
