Metadata-Version: 2.4
Name: factframe
Version: 0.1.0
Summary: PDF, docx and xlsx to markdown with a stable id on every block, table row and figure, and provenance back to the page. Ships a local MCP server.
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/Plinsboorg/factframe
Project-URL: Issues, https://github.com/Plinsboorg/factframe/issues
Keywords: mcp,pdf,docx,xlsx,markdown,citations,provenance
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Topic :: Text Processing :: Markup :: Markdown
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: local
Requires-Dist: pymupdf>=1.24; extra == "local"
Provides-Extra: azure
Requires-Dist: azure-ai-documentintelligence>=1.0.0; extra == "azure"
Requires-Dist: azure-core>=1.30; extra == "azure"
Requires-Dist: pypdf>=4.0; extra == "azure"
Provides-Extra: anthropic
Requires-Dist: anthropic>=1.0; extra == "anthropic"
Provides-Extra: openai
Requires-Dist: openai>=1.0; extra == "openai"
Provides-Extra: server
Requires-Dist: fastapi>=0.115; extra == "server"
Requires-Dist: uvicorn[standard]>=0.30; extra == "server"
Provides-Extra: mcp
Requires-Dist: mcp>=2.0; extra == "mcp"
Provides-Extra: xlsx
Requires-Dist: openpyxl>=3.1; extra == "xlsx"
Requires-Dist: defusedxml>=0.7; extra == "xlsx"
Provides-Extra: bridge
Requires-Dist: mcp>=2.0; extra == "bridge"
Requires-Dist: ollama>=0.3; extra == "bridge"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pymupdf>=1.24; extra == "dev"
Requires-Dist: httpx>=0.27; extra == "dev"
Requires-Dist: openpyxl>=3.1; extra == "dev"
Requires-Dist: defusedxml>=0.7; extra == "dev"
Dynamic: license-file

# FactFrame

PDF (and docx, xlsx) to markdown, where every block of text, every table row and every figure has
a stable id — and any id resolves back to a rectangle on a page of the original
PDF. An answer that cites an id is an answer a reader can check in one click.

```
pdf ──▶ provider ──▶ layout.json ──▶ annotate ──▶ document.md      (ids inline)
                                             └──▶ elements.json    (id → page + box)
                                             └──▶ manifest.json    (pages, counts)
```

```markdown
{{para:a1b2c3 bbox=1.00,2.35,7.50,3.10}}
The device operates from a single supply.

{{table:9f8e7d bbox=1.00,3.40,7.50,5.90}}
| {{row:0a2f13 table=9f8e7d idx=0}} Symbol | Min | Max |
| --- | --- | --- |
| {{row:44c1a0 table=9f8e7d idx=1}} VIH | 2.0 | 5.5 |
```

## Use it from your AI assistant (MCP)

FactFrame runs as a local [MCP](https://modelcontextprotocol.io) server: point
it at folders of PDF, docx and xlsx files and your assistant can convert, search,
read and cite them — with a link from every citation to the exact place in the
original. Everything happens on your machine; no document leaves it. Free for
use on your own computer.

One command, using [uv](https://docs.astral.sh/uv/):

```bash
uvx factframe-mcp install --root ~/Documents/contracts
```

That registers the server with the desktop clients it finds (Claude Desktop,
Cursor, Windsurf) and tells you what to restart. Repeat `--root` for more
folders. No uv yet?

```bash
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh && uvx factframe-mcp install
```

```powershell
# Windows (PowerShell)
irm https://astral.sh/uv/install.ps1 | iex; uvx factframe-mcp install
```

**Claude Code:**

```bash
claude mcp add factframe -- uvx factframe-mcp --root ~/Documents/contracts
```

Then ask your assistant to convert a folder and answer from it. Details and
options are under [MCP server](#mcp-server) below.

Not read: scanned PDFs with no text layer (there is no OCR in the local
version). Need your team's Google Drive or SharePoint connected, or a secured
deployment? Write to [dany@factframe.tech](mailto:dany@factframe.tech).

## Try it

```bash
pip install -e ".[local,dev]"

factframe convert your.pdf -o out/          # markdown + provenance index
factframe convert your.xlsx -o out/         # also .docx; needs the [xlsx] extra for workbooks
factframe show out/ 44c1a0                  # where did that element come from?
factframe ask out/ "What is the max supply voltage?"   # sourced answers (needs [openai])

# --llm anthropic for the Anthropic API instead of the default Azure OpenAI model:
factframe ask out/ "What is the max supply voltage?" --llm anthropic
```

### Web demo

Side-by-side highlight demo — rendered markdown on the left, the source PDF on
the right, click either to light up the other.

Install dependencies once:

```bash
cd web
npm install
```

The demo reads converted documents from `web/demo/public/<slug>/` (or, for a
folder of PDFs converted together, `web/demo/public/<folder>/<slug>/`) —
`document.md`, `elements.json`, `manifest.json`, plus the source PDF itself
as `source.pdf`. You can drop one in by hand:

```bash
factframe convert your.pdf -o web/demo/public/sample
cp your.pdf web/demo/public/sample/source.pdf
```

Then run it:

```bash
cd web
npm run demo
```

Vite serves it on `http://localhost:5173/` and opens it in a browser. Pick a
different port if that one is taken (`npm run demo -- --port 5174`). `/`
lists every converted document; each opens at `/doc/<slug>` (or
`/doc/<folder>/<slug>`), optionally with `?highlight=<id>` to land scrolled
and flashed on one element — a bookmarkable, shareable deep link.

You don't have to convert up front, either — the doc list has a **Convert &
view** button for a single PDF or .xlsx and a **Choose folder** picker to convert
a whole folder of them in one go. Both post to a dev-server-only `/api/convert`
endpoint (see `demo/vite.config.ts`), which shells out to `factframe convert`
(the same `local`/pymupdf provider, titled from the filename) into
`web/demo/public/<slugified-filename>/` (or, with a folder, one level
deeper). Needs `factframe` importable from the project's `.venv` (or on
`$PATH`) in whatever shell runs `npm run demo` — it only exists in the dev
server, not in `npm run build`.

### MCP server

A local, stdio-transport MCP server. Six tools:

| Tool | What it does |
|---|---|
| `convert_folder` | convert every PDF, docx and xlsx in a folder and its subfolders; unchanged files are skipped on the next run |
| `list_documents` | what is already converted |
| `search_documents` | keyword search across documents (or one folder / one document), returning element ids |
| `read_document` | a document's markdown, ids inline, a range of pages at a time |
| `get_element` | one element's text by id, to check a citation |
| `make_link` | a URL that opens the viewer scrolled and highlighted to one element |

Options, on the command line the client runs:

| Option | Meaning |
|---|---|
| `--root FOLDER` | a folder the server may read from; repeatable. Required: without one, `convert_folder` refuses to run |
| `--store DIR` | where conversions are written. Default `~/.factframe/store` |
| `--mirror` | also write a markdown copy next to each source file (`report.pdf.md`) |

`FACTFRAME_ALLOWED_ROOTS` (path-separator-separated), `FACTFRAME_PUBLIC_DIR`
and `FACTFRAME_MIRROR=1` are the same three as environment variables.

`uvx factframe-mcp install` writes the entry for you. To add it by hand, every
client reads the same shape, from different files (`claude_desktop_config.json`,
`.cursor/mcp.json`, `.vscode/mcp.json`, ...):

```json
{
  "mcpServers": {
    "factframe": {
      "command": "/absolute/path/to/uvx",
      "args": ["factframe-mcp", "--root", "/absolute/path/to/your/folder"]
    }
  }
}
```

Use absolute paths: a desktop app does not inherit your shell's `PATH`, and
some clients do not expand `~`. VS Code wants `"servers"` instead of
`"mcpServers"`.

**The viewer.** `make_link`'s URLs open a read-only viewer that the server
itself runs on `127.0.0.1` (port 47200, or a free one if that is taken),
started the first time a link is asked for. It lives as long as the server
does, so a link opens on this computer while the client is running. Set
`FACTFRAME_WEB_BASE_URL` to point links at a viewer hosted elsewhere instead.

**From a checkout.** An editable install (`pip install -e ".[mcp,local,xlsx]"`)
writes into `web/demo/public/`, so anything converted through the server shows
up in a running `npm run demo`; run `npm run viewer:build` in `web/` to give
the built-in viewer something to serve.

**Licensing.** FactFrame is Apache-2.0. The local PDF reader is
[PyMuPDF](https://pymupdf.readthedocs.io), which is AGPL-3.0 (or commercially
licensed by Artifex); that matters if you redistribute or host this as a
service, not for running it on your own machine.

## Configuration

Nothing is required to start. The default provider reads the PDF's own text
layer locally — no account, no key, no per-page cost. Credentials only unlock
the optional paths:

| Variable | Needed for | Notes |
|---|---|---|
| `AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT` | `--provider azure` | e.g. `https://<resource>.cognitiveservices.azure.com/` |
| `AZURE_DOCUMENT_INTELLIGENCE_KEY` | `--provider azure` | one key, or several separated by commas |
| `AZURE_DOCUMENT_INTELLIGENCE_PAGES_PER_CHUNK` | free-tier Azure keys | set to `2` — see below |
| `AZURE_OPENAI_ENDPOINT` | `factframe ask` (default `--llm openai`) | e.g. `https://<resource>.openai.azure.com/` |
| `AZURE_OPENAI_API_KEY` | `factframe ask` (default `--llm openai`) | |
| `AZURE_OPENAI_API_VERSION` | `factframe ask` (default `--llm openai`) | defaults to `2025-04-01-preview` (needed for the Responses API) |
| `ANTHROPIC_API_KEY` | `factframe ask --llm anthropic` | conversion and highlighting work without it |
| `FACTFRAME_AUTH_USER` | Web demo & API | Username for HTTP Basic Auth |
| `FACTFRAME_AUTH_PASS` | Web demo & API | Password for HTTP Basic Auth |
| `FACTFRAME_HOST` | `npm run demo` (Vite dev server) | Shell env var, not a `.env` entry — defaults to `127.0.0.1`; set to `0.0.0.0` (or a specific address) to opt into binding all interfaces, e.g. for a remote dev session over a forwarded port. Same idea as `HOST` for `npm run demo:serve` (`prod-server.mjs`), which already defaults to `127.0.0.1`. |

`factframe ask` defaults to `--llm openai`, a model deployed on Azure OpenAI
(`gpt-5.4-mini` by default) — `pip install factframe[openai]`. Pass `--llm
anthropic` to use the Anthropic API instead (`pip install
factframe[anthropic]`), and `--model` to override either adapter's default
model (for Azure OpenAI, this is the deployment name).

## Working with free-tier Azure keys

**The free (F0) tier reads only the first two pages of any document you submit.**
There is no error: pages three onward are simply absent from the result, so a
40-page datasheet silently converts to a two-page one. If you are on a free key,
this is the single most important thing to know about the provider.

The fix is to split the PDF into two-page pieces, analyse each one, and stitch
the results back into a single document:

```bash
export AZURE_DOCUMENT_INTELLIGENCE_ENDPOINT="https://<resource>.cognitiveservices.azure.com/"
export AZURE_DOCUMENT_INTELLIGENCE_KEY="key1,key2,key3"
export AZURE_DOCUMENT_INTELLIGENCE_PAGES_PER_CHUNK=2

factframe convert your.pdf -o out/ --provider azure
```

`AZURE_DOCUMENT_INTELLIGENCE_KEY` takes **several comma-separated keys**. Chunks
are spread across them round-robin, and a retry always moves to the next key —
so a key that has hit its rate or monthly page limit doesn't fail the same chunk
twice. Three free keys give you three times the throughput and three times the
monthly page allowance.

Everything lives in `providers/azure.py`: `split_pdf` cuts the PDF, `shift_pages`
moves each chunk's page numbers to where they really belong (a chunk always
comes back numbered from 1), and `_combine` reassembles the document.

### What stitching has to get right

Both of these are silent when wrong — nothing raises, you just get a document
that is subtly not the one you fed in. `tests/test_azure_chunking.py` pins both.

- **Page numbers must be shifted everywhere**, not just on the `pages` entries.
  Table cells and figure captions carry their own `boundingRegions`, and a region
  left pointing at page 1 puts every highlight from that chunk on the cover page.
- **Span offsets must be rebased.** Each element's `spans` index into its own
  chunk's `content` string. Concatenate the strings without shifting the offsets
  and most spans point at the wrong text — measured at ~81% wrong on a real
  document. (FactFrame itself reads neither `content` nor `spans`; it renders
  from the positioned elements. The stitch is correct anyway, because the layout
  document is the contract.)

### Caveats

- **On a paid tier, leave chunking off.** It is not an optimisation. The service
  costs roughly a fixed 84s per call plus ~0.11s per page, so on a 1380-page
  document five sequential chunks took ~573s against ~240s for a single call.
  Chunking wins only when a tier limit makes the single call impossible.
- **Anything a provider computes per document is computed per chunk.** A model
  that infers structure from statistics over the whole document sees only two
  pages at a time and may label a heading differently near a chunk boundary.
  Verified end to end on an 18-page datasheet processed two pages at a time: page
  numbering, tables, rows and figures all came out identical to a single-shot
  run, with two paragraphs reclassified as headings.

## Layout

| Path | What |
|---|---|
| `src/factframe/ids.py` | stable, content-derived element ids |
| `src/factframe/geometry.py` | page geometry, boxes, page-percentage conversion |
| `src/factframe/layout.py` | the layout document contract + id assignment |
| `src/factframe/blocks.py` | grouping paragraphs into citable blocks |
| `src/factframe/markdown.py` | markdown rendering with provenance sentinels |
| `src/factframe/elements.py` | the element index: id → kind, page, box, text |
| `src/factframe/provenance.py` | citation → page region; citation checking |
| `src/factframe/ask.py` | sourced answers from an LLM, citations validated |
| `src/factframe/providers/azure.py` | Azure Document Intelligence + free-tier chunking |
| `src/factframe/providers/pymupdf.py` | local, offline, free extraction |
| `src/factframe/docx_convert.py`, `xlsx_convert.py` | docx and workbooks → the same artifacts, no provider |
| `src/factframe/retrieval.py` | search over a workbook: filter, BM25, cell pinning |
| `src/factframe/webstore.py` | slug/path rules shared by the web demo and the MCP server |
| `src/factframe/mcp_server.py` | the MCP server — convert, list, search, read, link |
| `src/factframe/viewer.py` | the read-only viewer the MCP server runs for its links |
| `src/factframe/install.py` | `factframe-mcp install` — registers the server with a client |
| `packages/factframe-mcp/` | the launcher package behind `uvx factframe-mcp` |
| `web/src/` | markdown-it anchor plugin, resolver, highlight |
| `web/demo/` | the side-by-side viewer |

Every module's docstring explains not just what it does but why it does it that
way — the design decisions are the part that took the longest to get right.

## Design notes

- **Ids are content-derived, never random.** Re-converting the same PDF produces
  the same ids, so citations stored by an earlier run keep resolving.
- **A citation is a bare id and nothing else.** Ids are unique across kinds, so
  the kind is recovered from the index at read time. Every schema that stored the
  kind alongside the id eventually stored it wrong.
- **The provider is the only impure stage.** Everything after `layout.json` is a
  pure function, so re-rendering after a code change costs nothing and re-runs
  need no OCR.
- **Every citation is checked before it is trusted.** An invented citation is
  worse than a missing one: it presents as verified.

Run the tests with `pytest`.

Extracted from the FactFrame subsystem of datasheets.md and generalised to
arbitrary PDFs. Status: working end to end (Python pipeline, viewer, demo);
broader test coverage and `docs/` still to come.
