Metadata-Version: 2.4
Name: bibshelf
Version: 0.1.0
Summary: Shelve a paper or a book: fetch its bibtex, file its pdf
Keywords: bibtex,bibliography,doi,arxiv,isbn,reference-manager
Author: Arumoy Shome
Author-email: Arumoy Shome <contact@arumoy.me>
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: MacOS
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Text Processing :: Markup :: LaTeX
Requires-Python: >=3.11
Project-URL: Homepage, https://github.com/arumoy-shome/bibshelf
Project-URL: Issues, https://github.com/arumoy-shome/bibshelf/issues
Description-Content-Type: text/markdown

# bibshelf

Shelve a paper or a book: fetch its bibtex, file its pdf, name it something you
can find again.

```console
$ bs 10.1145/3411764.3445518 --pdf ~/Downloads/paper.pdf
bs: ~/Downloads/paper.pdf --> .../files/Sambasivan et al - 2021 - Everyone wants to do the model work, not the data work Data Cascades in High-Stakes AI.pdf
bs: proceed? [y/N]:
```

You get the renamed pdf, a `.bib` beside it, and the bibkey on your clipboard.

## Install

```console
uv tool install bibshelf     # or: pipx install bibshelf
```

The only optional dependency is `pdftotext` (from poppler), used to read an
identifier off a pdf. Without it that one feature asks you to type the
identifier instead; everything else works.

Copying to the clipboard uses `pbcopy` on macOS and `clip` on Windows, both of
which ship with the system. On Linux it wants one of `wl-copy`, `xclip` or
`xsel`; with none of them installed the entry is printed instead.

## Use

Give it a **DOI**, an **arXiv id**, or an **ISBN**:

```console
$ bs 10.1109/CAIN58948.2023.00034     # doi
$ bs 2211.09545                       # arxiv, new style
$ bs hep-th/9711200                   # arxiv, pre-2007
$ bs 978-0-262-03561-3                # isbn, hyphens optional
```

With `--pdf` it also files the pdf. Papers go to `files/`, books (anything with
an ISBN) go to `books/`, both under `~/Documents/references` — or wherever you
point it:

```console
$ export BIBSHELF_LIBRARY=~/work/library   # or, per run: bs --library ...
```

Both directories have to exist; bibshelf will not create them.

Leave the identifier out and it reads one off the pdf itself, then shows you the
entry it resolved before committing to anything:

```console
$ bs --pdf paper.pdf
bs: paper.pdf says 10.1007/s10664-023-10291-1

@article{morovati2023bugs,
  author = {Morovati, Mohammad Mehdi and Nikanjam, Amin and Khomh, Foutse},
  title = {Bugs in Machine Learning-Based Systems: A Faultload Benchmark},
  ...
}

bs: does this match the pdf? [Y/n]:
```

If the pdf mentions several identifiers — conference papers often carry both
their own doi and their proceedings' — it lists them and lets you pick. Say no
to one and it offers the next.

| flag | |
| --- | --- |
| `-p, --pdf PATH` | file this pdf alongside the entry |
| `-l, --library PATH` | where the library lives, overriding `$BIBSHELF_LIBRARY` |
| `-f, --force` | don't ask before moving anything |
| `-x, --to-clipboard` | copy the whole entry rather than just the bibkey |
| `-V, --version` | print the version |

## How files are named

```
Sambasivan et al - 2021 - Everyone wants to do the model work Data Cascades in High-Stakes AI.pdf
[  first creator  ] [year] [                     title                                       ]
```

One author is `Fowler`, two are `Aldiabat and Le Navenec`, three or more are
`Nahar et al`. Missing pieces drop out, so an undated book is just
`Barocas - Fairness and Machine Learning.pdf`.

Colons and double quotes are removed, accents are folded (`Géron` becomes
`Geron`), and the whole name is capped at the 255 bytes a filename allows,
cut on a character boundary. Accents survive in the bibtex `author` field —
they are only stripped from filenames and bibkeys, which latex is fussy about.

## Where the metadata comes from

DOIs and arXiv ids both resolve through `doi.org` content negotiation, which
returns CSL-JSON for Crossref and DataCite alike; an arXiv id becomes the doi
arXiv minted for it (`10.48550/arXiv.<id>`). ISBNs go to Open Library, whose
coverage is thinner — expect to eyeball book entries more than paper ones.

Entries are rendered directly from CSL. Titles are emitted verbatim, without
brace protection or title-casing, which means a style like `plain.bst` may
re-case them.

## Licence

MIT.
