Metadata-Version: 2.5
Name: pydfdoi
Version: 0.1.1
Summary: Extract DOI values from PDFs, write DOI metadata, and rename papers from Crossref citation data.
Project-URL: Homepage, https://github.com/Yihtsy/pydfdoi
Project-URL: Repository, https://github.com/Yihtsy/pydfdoi
Project-URL: Issues, https://github.com/Yihtsy/pydfdoi/issues
Project-URL: Changelog, https://github.com/Yihtsy/pydfdoi/blob/main/CHANGELOG.md
Author-email: Yihtsy <yihtsy@outlook.com>
License: MIT License
        
        Copyright (c) 2026 Yihtsy
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: crossref,doi,metadata,pdf,rename
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Utilities
Requires-Python: >=3.14
Requires-Dist: pikepdf>=10
Requires-Dist: pypdf>=5
Requires-Dist: requests>=2.31
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: twine>=5; extra == 'dev'
Description-Content-Type: text/markdown

# pydfdoi

`pydfdoi` is a small Python package for DOI-aware PDF cleanup. It extracts DOI
values from PDF metadata or PDF text, writes DOI-related PDF metadata, and can
rename academic PDFs from Crossref citation data.

Author: Yihtsy <yihtsy@outlook.com>

## Features

- Extract DOI values from existing PDF metadata and the first pages of text.
- Write `/doi` and `/doiURL` into PDF metadata.
- Set `/Title` to the file name for metadata-only cleanup.
- Configure PDFs to open on the first page.
- Ask PDF viewers to show the file name in the title bar.
- Query Crossref metadata for DOI records.
- Rename journal article PDFs as `Author et al. Year Title.pdf`.
- Align journal article PDF page labels with Crossref page metadata.
- Capture journal DOI metadata, book DOI metadata, or general DOI metadata.
- Look up common publisher abbreviations and DOI prefixes.
- Install optional Windows context-menu and SendTo shortcuts.

## Package Name

The project uses one consistent package name:

- PyPI distribution: `pydfdoi`
- Python import: `pydfdoi`
- Console command: `pydfdoi`
- Compatibility command: `pdfdoi`

This avoids the common mismatch where a package is installed with a hyphenated
name but imported with an underscore.

## Installation

After the package is published to PyPI:

```powershell
python -m pip install pydfdoi
```

For local development from a cloned repository:

```powershell
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -U pip
python -m pip install -e ".[dev]"
```

Editable installation means changes under `src\pydfdoi` are immediately used by
the command line and tests. Do not develop directly inside Python's
`site-packages`; keep the source in a normal project folder and use editable
installation.

## Command Line

Update DOI metadata only:

```powershell
pydfdoi --mode metadata-only "D:\path\paper.pdf"
```

Rename one or more PDFs from Crossref citation metadata:

```powershell
pydfdoi --mode rename "D:\path\paper1.pdf" "D:\path\paper2.pdf"
```

Scan more pages for DOI text:

```powershell
pydfdoi --mode metadata-only --max-pages 5 "D:\path\paper.pdf"
```

The command-line default is 20 pages, which is friendlier to ebooks whose DOI
often appears on the title or copyright pages rather than the first page.

Preview actions without writing changes:

```powershell
pydfdoi --mode rename --dry-run "D:\path\paper.pdf"
```

Print JSON output:

```powershell
pydfdoi --mode metadata-only --json "D:\path\paper.pdf"
```

Align PDF page labels with Crossref journal article page metadata:

```powershell
pydfdoi --mode page-labels "D:\path\article.pdf"
```

If the PDF has extra non-article pages, they are labeled with `skip-` by
default. Extra pages are assumed to be at the front unless you choose another
placement:

```powershell
pydfdoi --mode page-labels --extra-placement front "D:\path\article.pdf"
pydfdoi --mode page-labels --extra-placement back "D:\path\article.pdf"
pydfdoi --mode page-labels --extra-placement split "D:\path\article.pdf"
```

You can customize the special label prefix:

```powershell
pydfdoi --mode page-labels --skipped-label-prefix "nonarticle-" "D:\path\article.pdf"
```

If no file is passed, the command opens a file picker.

## Python API

### Metadata-only cleanup

Use `capture_and_write_doi_metadata()` when you only want to capture DOI
information and update the PDF metadata.

```python
from pydfdoi import capture_and_write_doi_metadata

doi = capture_and_write_doi_metadata("paper.pdf")
```

This function:

- extracts the DOI from the PDF;
- writes `/doi`;
- writes `/doiURL`;
- sets `/Title` to the file name without `.pdf`;
- configures the PDF to open on the first page;
- does not query Crossref;
- does not rename the file.

### Rename by citation

```python
from pathlib import Path

from pydfdoi import extract_doi, rename_by_citation

path = Path("paper.pdf")
doi = extract_doi(path)
if doi:
    renamed_path = rename_by_citation(path, doi)
```

### Journal, book, and general DOI capture

```python
from pydfdoi import capture_book_doi, capture_doi_metadata, capture_journal_doi

journal = capture_journal_doi("article.pdf")
book = capture_book_doi("book.pdf")
any_work = capture_doi_metadata("unknown.pdf")
```

- `capture_journal_doi()` accepts Crossref journal-like work types.
- `capture_book_doi()` accepts Crossref book-like work types.
- `capture_doi_metadata()` is the general DOI metadata entry point.

`capture_book_doi()` scans more pages by default than the journal helper because
books often put DOI information on title, copyright, or series pages rather than
the first page. Its default is 20 pages:

```python
book = capture_book_doi("book.pdf")
book = capture_book_doi("book.pdf", max_pages=30)
book = capture_book_doi("book.pdf", strict_type=True)
```

When `strict_type=False`, which is the default, `capture_book_doi()` can still
accept a DOI if Crossref returns an unknown work type but the page-count
heuristic says the PDF is book-like. If Crossref clearly identifies the work as
journal-like, it is still rejected by the book helper.

The general entry point can also infer a coarse work category:

```python
result = capture_doi_metadata("unknown.pdf", article_page_limit=80)

print(result.inferred_kind)           # "journal", "book", or None
print(result.classification_source)   # "page-count", "crossref", or None
print(result.page_count)
```

By default, `capture_doi_metadata()` first makes a simple page-count guess:
PDFs with 80 pages or fewer are treated as journal-like, and longer PDFs are
treated as book-like. If Crossref returns a known work type, the Crossref type
overrides the page-count guess.

### Publisher helpers

```python
from pydfdoi import PUBLISHERS, publisher_for_doi, publisher_for_name

publisher = publisher_for_doi("10.1038/s41586-020-2649-2")
publisher = publisher_for_name("John Wiley & Sons")
```

The publisher table is a curated starter list of common publishers, DOI
prefixes, abbreviations, and aliases. It is intentionally not exhaustive:
DOI prefixes are assigned to registrants and may remain in use after publisher
mergers, platform changes, or imprint transfers.

### Journal article page labels

Use `align_journal_article_page_labels()` to align PDF page labels with Crossref
article page metadata.

```python
from pydfdoi import align_journal_article_page_labels

plan = align_journal_article_page_labels(
    "article.pdf",
    extra_placement="front",
    skipped_label_prefix="skip-",
)
```

For a Crossref page range such as `10-12`, a four-page PDF is labeled as:

```text
skip-1, 10, 11, 12
```

The function can also use Crossref `article-number` when a conventional page
range is absent. For article-number-only records, labels look like `e12345-1`,
`e12345-2`, and so on.

Because Crossref usually does not say whether extra PDF pages are at the front
or the back, `extra_placement` is explicit:

- `front`: extra pages are before the article content.
- `back`: extra pages are after the article content.
- `split`: extra pages are split between front and back.

## Windows Context Menu

The repository includes compatibility scripts for Windows Explorer workflows.

Run PowerShell in the project folder:

```powershell
.\install_context_menu.ps1
```

This adds two PDF right-click entries:

- `PDF DOI: rename by citation`
- `PDF DOI: metadata only`

It also installs SendTo shortcuts, which are usually the most reliable route
for processing multiple selected PDFs in Windows Explorer.

To uninstall:

```powershell
.\install_context_menu.ps1 -Uninstall
```

`pdf_doi_tool.py` is kept as a compatibility wrapper for these scripts and for
older local workflows.

## Development

Project layout:

```text
src/
  pydfdoi/
    __init__.py
    citation.py
    cli.py
    crossref.py
    doi.py
    metadata.py
    page_labels.py
    pdf.py
    processing.py
    publishers.py
tests/
```

Run tests:

```powershell
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m pytest
```

Build local distributions:

```powershell
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m build
```

Check built distributions before publishing:

```powershell
C:\Users\yezi_\AppData\Local\Programs\Python\Python314\python.exe -m twine check dist/*
```

Publisher-specific observations and heuristic notes live in
[docs/development-notes.md](docs/development-notes.md).

## Release Notes

Current version: `0.1.1`

See [CHANGELOG.md](CHANGELOG.md) for the full version history.

## Publishing

Before publishing a new release:

1. Update `version` in `pyproject.toml`.
2. Update `__version__` in `src/pydfdoi/__init__.py`.
3. Add a new entry to `CHANGELOG.md`.
4. Run `python -m pytest`.
5. Run `python -m build`.
6. Run `python -m twine check dist/*`.
7. Tag the release in Git.
8. Upload to PyPI with `python -m twine upload dist/*`.

For the first PyPI upload, create an account-wide API token on PyPI and run the
local helper:

```powershell
.\scripts\publish_pypi.ps1 -SkipBuild
```

The helper prompts for the token without echoing it, sets Twine credentials only
for that process, uploads the built `0.1.0` distributions, and clears the
temporary environment variables afterward.

## License

`pydfdoi` is released under the MIT License. See [LICENSE](LICENSE).
