Metadata-Version: 2.4
Name: zotkit
Version: 0.5.2
Summary: Headless Zotero library management: Web API CRUD plus direct WebDAV attachment upload/download — no desktop app required
Author: Shawn
License-Expression: MIT
Project-URL: Repository, https://github.com/oldantique/zotkit
Project-URL: Issues, https://github.com/oldantique/zotkit/issues
Keywords: zotero,reference-manager,webdav,bibliography,cli,headless
Classifier: Programming Language :: Python :: 3
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pyzotero
Requires-Dist: httpx
Dynamic: license-file

# zotkit

[![PyPI](https://img.shields.io/pypi/v/zotkit)](https://pypi.org/project/zotkit/)
[![Python](https://img.shields.io/pypi/pyversions/zotkit)](https://pypi.org/project/zotkit/)
[![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)

**Headless Zotero library management — no desktop app required.**

**English** | [简体中文](README.zh-CN.md)

"Headless" simply means zotkit never needs the Zotero app (or any window) open: it is a
Python library + CLI that talks straight to the
[Zotero Web API](https://www.zotero.org/support/dev/web_api/v3/start), so you can
search, create, tag, and organize items from any terminal — macOS, Windows, or Linux,
your laptop or a remote server. If your attachments sync to a **personal WebDAV
server**, zotkit can **upload and download the files themselves** by speaking
Zotero's WebDAV storage format directly — a capability the Web API itself does not
provide. The format is documented in [docs/webdav-format.md](docs/webdav-format.md).

Built for servers, scripts, and **LLM agents**: every write is dry-run by default,
batched, and version-checked, and you can define a tag taxonomy that is *enforced in
code* so an agent (or a tired human) can't pollute your library with inconsistent tags.

## Why zotkit

| | Desktop app | Other CLI/MCP tools | zotkit |
|---|---|---|---|
| Works headless (server, SSH, CI) | ❌ | ✅ read-mostly | ✅ |
| Write items/tags/collections | ✅ | ⚠️ usually needs the desktop app running | ✅ |
| Attachment files (Zotero Storage) | ✅ | ⚠️ some | ✅ upload + download |
| Attachment files on **WebDAV** | ✅ | ⚠️ download at best | ✅ **upload + download** |
| Tag conventions enforced in code | ❌ | ❌ | ✅ optional `conventions.toml` |

## The zotkit family

Two independent, complementary projects — use either on its own, or both together:

| | What it is |
|---|---|
| **zotkit** (this repo) | Headless Python CLI + library. Runs anywhere, needs no Zotero app. |
| **[zotkit-reader](https://github.com/oldantique/zotkit-reader)** | A Zotero Reader sidebar that embeds a Codex/Claude agent beside the PDF you're reading. Maintained by [@ChanceSiyuan](https://github.com/ChanceSiyuan). |

They share a name and a philosophy, not a codebase: zotkit-reader can call zotkit over MCP
for library-wide operations, but neither requires the other to be installed.

## Install

Pure Python (3.11+), no platform-specific bits — the same package works on macOS,
Windows, and Linux:

```bash
pipx install zotkit        # or: uv tool install zotkit / pip install zotkit
uvx zotkit --help          # …or try it without installing anything
```

## Configure

Copy [`.env.example`](.env.example) to `./.env`, `~/.config/zotkit/env`, or any path in
`$ZOTKIT_ENV`, and fill in:

- **Zotero Web API**: create a key (with write access) at
  <https://www.zotero.org/settings/keys> — your numeric `ZOTERO_LIBRARY_ID` is shown on
  the same page.
- **WebDAV** (only for `attach`/`fetch`): copy the exact values from the Zotero desktop
  app on any of your machines — **Settings → Sync → File Syncing** — and append
  `/zotero/` to the URL (the desktop does this implicitly).
- **Using Zotero Storage instead of WebDAV?** Just leave the `WEBDAV_*` lines out —
  `attach`/`fetch` automatically use Zotero Storage through the Web API's upload/download
  endpoints instead. The storage mode is detected from your `.env`, nothing to configure.

After filling it in, run **`zotkit doctor`** — it validates the config file, API
access, and attachment storage, and tells you exactly what to fix if anything fails.

Optionally, copy [`conventions.example.toml`](conventions.example.toml) to
`conventions.toml` next to your `.env` to define a namespaced tag taxonomy
(`field:physics`, `status:to-read`, …). With it in place, `zotkit create` / `zotkit tag`
**reject** violations; without it, tags are unrestricted.

## Quickstart

```bash
zotkit find --title "boson sampling"        # search by title/tag/collection
zotkit find --tag status:to-read

zotkit create --arxiv 2401.12345            # fetch arXiv metadata, dry-run preview
zotkit create --arxiv 2401.12345 --apply --tags field:ai   # create + download & attach the PDF
zotkit create --arxiv 2401.12345 1706.03762 math/0211159 --apply   # batch: one metadata
                                            #   request, PDFs politely spaced ≥3 s apart
zotkit create --doi 10.1038/nature14539     # fetch CrossRef metadata by DOI (no PDF —
                                            #   usually paywalled; attach manually after)

zotkit create --file papers.json            # batch from JSON: dry-run preview
zotkit create --file papers.json --apply    # create (dedups by DOI/title)
zotkit attach --from papers.created.json --all   # upload the PDFs (WebDAV or Zotero
                                                 #   Storage — auto-detected from .env)

zotkit enrich --key AB12CD34 EF56GH78       # fill missing fields on existing items
                                            #   in place (dry-run; --apply to write)

zotkit attach --key AB12CD34 --pdf paper.pdf     # single attach
zotkit fetch --key AB12CD34 --out downloads      # download attachments (same auto-detection)

zotkit tag AB12CD34 topic:qaoa prio:high    # validated against conventions.toml
zotkit status AB12CD34 read                 # replaces the status: tag
zotkit move AB12CD34 "Algorithms"           # or "Parent :: Child"; --add keeps old home

zotkit backup                               # full JSON snapshot -> backups/
zotkit lint field:physics topic:new-idea    # offline tag check
```

`--arxiv` takes ids or abs/pdf URLs (several, space- or comma-separated) and maps
the full record (all authors, abstract, date, DOI); `--doi` maps CrossRef records
(journal articles, conference papers, books, chapters, …) and refuses to guess on
CrossRef types it doesn't know. Both accept `--collection` and `--tags`, and
`--no-pdf` skips the arXiv PDFs. Rate limiting is built into the request layer —
batches use one arXiv metadata request and space PDF downloads per arXiv's terms
of use, so callers (humans or agents) never pace themselves. A bad id fails
alone, not the batch; the exit code is non-zero only if something failed.
Fetching metadata from arbitrary web pages is out of scope (zotkit stays a
daemon-free CLI — no translation-server). CrossRef requests identify themselves
to the polite pool with a contact address; that should be reachable, so if you
distribute a tool built on zotkit, set your own via `ZOTKIT_MAILTO`.

**Version of record**: when arXiv reports a *journal* DOI (the paper was formally
published), `--arxiv` builds the journal record from CrossRef instead of a
`preprint`: proper item
type, venue, volume/pages, formal date. The arXiv identity is kept — `arXiv: <id>`
goes in Extra, the `url` stays the open-access abs page (the journal link lives in
the DOI field), the arXiv abstract fills in when CrossRef has none, and the PDF
still comes from arXiv. If the CrossRef lookup fails, the item falls back to the
preprint record with a warning rather than failing. A *repository* DOI is never
mistaken for a journal one: DOIs under the preprint servers' own prefixes
(arXiv `10.48550`, SSRN `10.2139`, bioRxiv/medRxiv `10.1101`, Research Square,
OSF, ChemRxiv, TechRxiv, Preprints.org) don't trigger the upgrade, and neither
does a CrossRef record that turns out to be a repository posting in disguise
(SSRN registers working papers as articles in a fake "SSRN Electronic
Journal") — those stay preprints, with the reason stated. The same predicate
gates `enrich --rebuild-record`. Items whose abstract zotkit
wrote also carry an `abstract-source: arxiv|crossref` line in Extra, naming where
it actually came from.

### Enriching existing items

`zotkit enrich --key K [K …]` completes incomplete items — missing abstracts,
DOIs, truncated author lists — from the same arXiv/CrossRef sources, **in
place**. The item key never changes (keys are the stable handle downstream
tooling references; delete-and-recreate is not an option), and writes carry the
item version, so a concurrent edit fails that item loudly instead of clobbering.

The merge is deliberately conservative: only empty fields are filled; creators
are extended only when the current list is a same-order prefix of the
authoritative one (the classic truncated-list case — any other difference is
reported, not touched); tags, collections, relations, and attachments are never
modified; Extra only gains lines. Abstracts zotkit writes are stamped
`abstract-source: …` in Extra — a stamp with no abstract means someone
deliberately removed it, and enrich will not re-add it (reported as NEEDS
OWNER). Items with no DOI or arXiv id are reported as needs-identifier.

`--rebuild-record` additionally upgrades a preprint whose paper has since been
published (journal DOI present) to the journal record **in the same item**:
itemType, published title/date/venue/volume/pages — with `arXiv: <id>` appended
to Extra and attachments untouched.

Item JSON for `zotkit create --file` (a list, one object per reference):

```json
[{"itemType": "journalArticle", "title": "…",
  "creators": [{"creatorType": "author", "firstName": "A", "lastName": "B"}],
  "date": "2024", "publicationTitle": "…", "DOI": "10.x/y",
  "tags": ["field:physics", "status:to-read"],
  "collection": "Algorithms", "file_path": "/abs/path/paper.pdf"}]
```

## From Python

```python
from zotkit import Zot

z = Zot()                                   # reads .env automatically
z.find(tag="status:to-read")
z.create_items([...])                       # dedup + convention checks
z.attach("AB12CD34", "paper.pdf")           # PDF -> WebDAV / Zotero Storage
z.fetch("AB12CD34", "downloads")
z.set_status("AB12CD34", "read")
z.backup()
```

`z.z` is the underlying [pyzotero](https://github.com/urschrei/pyzotero) client for
anything not wrapped.

## Using zotkit with AI agents

zotkit is designed to be driven by coding agents (Claude Code and similar): dry-run
defaults, code-enforced tag conventions, and a ready-made **Claude Code skill** in
[`skills/zotkit/`](skills/zotkit/SKILL.md) — copy it to `~/.claude/skills/zotkit/` and
any Claude session can search, file, and attach papers for you while respecting your
taxonomy.

Want to clean up a messy library, not just maintain one? The battle-tested method —
taxonomy design, parallel read-only analysis, serial reviewed writes — is written up in
[`docs/organizing-with-agents.md`](docs/organizing-with-agents.md).

```bash
mkdir -p ~/.claude/skills && cp -r skills/zotkit ~/.claude/skills/
```

## Safety model

- `create` is **dry-run by default**; `--apply` to execute.
- Writes go through fetch→modify→update (carries the item version, so concurrent edits
  fail loudly with 412 instead of clobbering), in batches of ≤ 50.
- `zotkit backup` snapshots every item, collection, tag, and membership to one JSON
  file — run it before bulk operations.
- Remember: writes propagate to zotero.org and **all your synced devices**.

## How WebDAV attachments work

(With Zotero Storage, zotkit simply uses the Web API's official file endpoints — this
section is about the WebDAV mode.) Zotero's WebDAV storage format is undocumented but
simple: each attachment item `K` is stored as `K.zip` (the file, zipped) plus `K.prop`
(its md5 + mtime). zotkit creates the attachment item via the Web API and PUTs both
objects directly — after which every desktop client syncs the file down normally.
Details in [`docs/webdav-format.md`](docs/webdav-format.md).

The format was determined by interoperability inspection of the author's own library.
This project is not affiliated with or endorsed by Zotero.

## Limits & roadmap

- `find` currently lists the library client-side — instant for hundreds of items,
  sluggish for many thousands. Server-side search is planned.
- Group libraries should work for item operations (untested); WebDAV file sync is
  personal-libraries-only (a Zotero limitation).
- `--doi`/`--arxiv` import covers arXiv + CrossRef; DataCite-only DOIs and
  arbitrary-URL scraping (translation-server territory) are out of scope.
- Planned: an MCP server wrapper, server-side search.

## License

[MIT](LICENSE). If you build on the WebDAV implementation, a link back is appreciated.
