Metadata-Version: 2.4
Name: watfile
Version: 0.2.1
Summary: Classify documents with the TypeSafe Jev decision model and sort them into folders
Keywords: classification,files,llm,typesafe,cli
Author: Michael Hunger
Author-email: Michael Hunger <github@jexp.de>
License-Expression: MIT
Classifier: Environment :: Console
Classifier: Intended Audience :: End Users/Desktop
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Utilities
Requires-Dist: liteparse>=2.14.6
Requires-Dist: tiktoken>=0.14.0
Requires-Dist: typesafe-sdk>=0.7.0
Requires-Dist: laya-mlx>=0.1.0 ; extra == 'laya'
Requires-Dist: laya-coreml>=0.1.0 ; extra == 'laya-coreml'
Requires-Dist: laya>=0.3.4 ; extra == 'laya-torch'
Requires-Python: >=3.12
Project-URL: Repository, https://github.com/jexp/watfile
Provides-Extra: laya
Provides-Extra: laya-coreml
Provides-Extra: laya-torch
Description-Content-Type: text/markdown

# watfile

Classify files with a decision-making AI model and sort them into category folders.

watfile sends each document's text (title/abstract-grade extract) to a
**TypeSafe AI Jev** (System One) [Choice](https://docs.typesafe.ai/primitives/choice)
question, gets back a typed answer with a selected category, per-category
probabilities and confidence, then moves the file into the matching folder.
A `Classifier` abstraction keeps the backend pluggable — local MLX (laya) and
other backends slot in later.

## Install

Requires Python 3.12+ and [uv](https://docs.astral.sh/uv/).

### From PyPI

```sh
# one-off run, no install
uvx watfile --help

# persistent CLI on your PATH
uv tool install watfile
watfile --help
```

Updates:

```sh
uv tool upgrade watfile      # or: watfile --self-update
uvx watfile@latest ...       # one-off runs always fetch the newest version
```

### Configuration

watfile resolves its TypeSafe API key (create one at <https://console.typesafe.ai/>)
with this precedence — first match wins:

1. `TYPESAFE_API_KEY` environment variable
2. `.env` file in the current directory (gitignored; `TYPESAFE_API_KEY=...`)
3. `~/.config/watfile/config.toml` (`api_key = "..."`, also `base_url`, `model`;
   `$WATFILE_CONFIG` or `$XDG_CONFIG_HOME` can relocate it)

```sh
export TYPESAFE_API_KEY=...        # option 1
echo 'TYPESAFE_API_KEY=...' > .env # option 2
cat > ~/.config/watfile/config.toml <<'EOF'   # option 3
api_key = "..."
EOF
```

## Usage

Point watfile at files or folders, and either give a comma-separated category
list (`-c`) or a target folder whose subfolders are the categories (`-d`):

```sh
# explicit categories, files moved into ./sorted/<category>/
uv run watfile ~/Downloads/invoice.pdf -c invoice,donation,apartment

# folder input, recursive; categories = existing subfolders of -d
mkdir -p ~/docs/{invoice,donation,apartment}
uv run watfile ~/Downloads -r -d ~/docs

# preview without touching anything
uv run watfile ~/Downloads -r -d ~/docs -n

# actually move the files (default is symlinking into the category folders)
uv run watfile ~/Downloads -r -d ~/docs -m

# copy instead
uv run watfile ~/Downloads -r -d ~/docs --copy

# custom output root with -c
uv run watfile *.pdf -c computerscience,biology -o ~/sorted
```

Output per file:

```
bill.pdf: invoice (conf 0.94) -> symlink to ~/docs/invoice/bill.pdf
```

Files that can't be classified (unsupported extension, no extractable text) are
skipped with a warning; name collisions get a `_1`, `_2`… suffix.

### Supported inputs

- **Text formats** (read directly): `.txt .md .markdown .rst .log .csv .json`
- **PDF** (via [liteparse](https://github.com/run-llama/liteparse)): only the
  first 2 pages are parsed, OCR disabled — enough for classification, ~1000x
  faster than a full parse. Scanned/image-only PDFs are skipped.

### Backends

- **`jev` (default)** — [TypeSafe AI](https://docs.typesafe.ai) Jev, cloud API.
  Needs `TYPESAFE_API_KEY`. Highest accuracy (4/4 on the arXiv fixtures).
  **Batches automatically**: documents are packed into one `system_one` call
  (~256 tokens each, up to ~100 files per call in the 30k-token window), so
  classifying a folder costs one API call, not one per file.
- **`laya`** — local typed-decision model from the
  [Laya](https://github.com/NandhaKishorM/laya) family; **runs on any
  platform**. The runtime is auto-detected from what's installed
  (`LAYA_RUNTIME` overrides):

  | extra | runtime | where | speed |
  |---|---|---|---|
  | `pip install watfile[laya]` | MLX (GPU) | Apple Silicon | ~13ms/decision |
  | `pip install watfile[laya-coreml]` | Core ML (ANE) | Apple Silicon | ~5ms, 2.8× lower energy |
  | `pip install watfile[laya-torch]` | PyTorch (CPU/GPU) | **any OS** | ~45–450ms (CPU) |

  No API key needed; checkpoints download once and then run offline. Default
  checkpoints: multilingual where available (torch/coreml → handles non-English
  documents out of the box). Because laya's context window is small
  (512–1024 tokens), the backend defaults to **adaptive multi-chunk
  classification**: the extract is split into ~200-token chunks; chunk 1
  decides if its probability is decisive (≥0.5), otherwise further chunks are
  classified and probabilities aggregated until the decision is decisive
  (max 10). On the arXiv fixtures: 3/4 (Jev 4/4).

```sh
watfile ~/Downloads -r -d ~/docs --backend laya
# checkpoint via config or env:
#   ~/.config/watfile/config.toml -> laya_model = "..."
#   or LAYA_MODEL=... / LAYA_RUNTIME=torch|mlx|coreml
```

### Options (main)

`watfile --help` shows only these:

```
usage: watfile [-h] [-r] (-c CATEGORIES | -d DIRECTORY) [-o OUTPUT]
               [--backend {jev,laya}] [-n] [-m | --copy | --symlink]
               [--help-all]
               inputs [inputs ...]

main options:
  -h, --help            show this help message and exit
  --help-all            show advanced options too
  -r, --recursive       recurse into folder inputs
  -v, --version         print version and exit
  -c CATEGORIES         comma-separated categories
  -d DIRECTORY          target folder whose existing subfolders are the categories
  -o OUTPUT             output root for sorted files (default: same as -d, or ./sorted with -c)
  --backend {jev,laya}  classifier backend (default: jev)
  -n, --dry-run         print decisions without placing files
  -m, --move            move files into the category folder (default: symlink)
  --copy                copy files instead of symlinking
  --symlink             create symlinks in category folders (default)
  --self-update         update watfile in place (uv tool / pipx aware)
```

### Advanced options

Shown by `watfile --help-all`:

```
  --batch N             cap files per API call (default: automatic — jev packs
                        everything that fits the 30k-token window, ~100 docs;
                        laya doesn't batch)
  --no-batch            disable batching, one API call per file (debugging)
  --chunk-tokens N      per-document token budget (default: backend-specific)
  --chunks N            split each document into N chunks, aggregate
                        probabilities; 0 = adaptive. Default: 0 for laya,
                        1 for jev
```

Batching example (jev batches by default; the flag just caps batch size):

```sh
# 100 files: ~2 API calls instead of 100 (256 tokens/doc, 30k window)
uv run watfile ~/Downloads -r -d ~/docs

# cap batch size, e.g. to keep batches small for debugging
uv run watfile ~/Downloads -r -d ~/docs --batch 25
```

## Development

### From source

```sh
git clone <repo> && cd watfile
uv sync            # create venv + install deps (typesafe-sdk, liteparse, laya-mlx)
uv run watfile --help

# or install the local checkout as a tool
uv tool install --from . watfile
```

### Tests

```sh
uv sync
uv run pytest              # unit tests; live API tests skip without TYPESAFE_API_KEY
```

`tests/fixture/` contains 4 real arXiv PDFs with ground-truth categories
(derived from their arXiv subject tags) used by the integration tests.

## Roadmap

- ~~`laya` local backend (MLX via OpenAI-compatible HTTP)~~ done — native laya-mlx
- ~~batching: classify 25/50/100 files in a single API call~~ done — jev batches
  automatically into the 30k-token window (default on, `--no-batch` to disable)
