Metadata-Version: 2.5
Name: clapback-cli
Version: 0.2.0
Summary: Search your own music by description, and find duplicates across formats and masters
License-Expression: MIT
License-File: LICENSE
Requires-Python: >=3.11
Requires-Dist: clapback-client<0.5,>=0.4.0
Requires-Dist: clapback-embed<0.2,>=0.1.0
Requires-Dist: numpy>=1.24.0
Provides-Extra: dev
Requires-Dist: pytest>=7.4.0; extra == 'dev'
Requires-Dist: ruff>=0.1.0; extra == 'dev'
Description-Content-Type: text/markdown

# clapback

Search your own music by description, and find duplicates across formats and masters.

```bash
pip install clapback-cli

clapback index ~/Music
clapback search "dreamy ambient with piano"
clapback duplicates
```

## What it does

**Search by description.** CLAP puts audio and text in one space, so "something
slow with brushed drums" is a query rather than a keyword match against filenames
you may never have typed.

**Find near-duplicates.** Two rips of one recording measure 0.9972–0.9995 under
this pipeline; genuinely different music sits far below. That gap is what makes
duplicate detection across formats and masters work — a FLAC and a V0 of the same
master are obvious, and so is the same recording on two different releases.

Both run against your own files, offline. There is no account, no key, and
nothing is sent anywhere.

**Contribute, if you want to.** Opt-in and off unless you type it:

```bash
clapback contribute --dry-run   # say what would be sent, send nothing
clapback contribute
```

This sends the vectors — never your audio, never filenames, never your library's
contents. A recording is identified by the SHA256 of its AcoustID fingerprint,
which is one-way, and this tool sends no recording id. Be clear about what that
does and does not hide: the corpus cannot recover a title from the hash, but if
another contributor has already named that same hash — 89.6% of rows are named —
the corpus knows which recording your row is. Everything sent is dedicated to the
public domain under CC0 1.0, like every other row in the corpus, and may be
republished in its public exports.

The whole library is looked up before anything is offered — a hundred tracks a
request — so re-running contributes only what is new. That is not politeness about
bandwidth: a repeat submission is recorded as agreement, and one install agreeing
with itself would corrupt the one measurement the commons exists to make.
Contributions go out a hundred at a time too, every guarantee per row, and a
refused row is that row's result rather than the run's.

Contributing needs `chromaprint`, and only contributing does:

```bash
brew install chromaprint     # or: apt install libchromaprint-tools
pip install pyacoustid
```

Without it, indexing, search and duplicates work exactly as well.

## Why `clapback-cli` and not `clapback`

The bare name on PyPI belongs to an unrelated package from 2018 that adds clap
emojis to sentences. The distribution is therefore `clapback-cli`, matching
`clapback-embed`; the command you type is still `clapback`.

## What it needs

`clapback-embed`, which arrives with it, and the ONNX encoders it runs on. Those
are **614 MB and not bundled** — a package that downloaded them on install would
be lying about its size. Export them once:

```bash
pip install 'clapback-embed[export]'
python -m clapback_embed.scripts.export_models --out ~/.cache/clapback/models
```

Or point `CLAPBACK_MODEL_DIR` at them if you already have them.

## Where things are kept

`~/.clapback/` — a `vectors.npy` and an `index.json`, both yours. Deleting the
directory loses nothing but the time to rebuild it.

If you contribute, `index.json` also holds a `client_id`: a random UUID minted the
first time you contribute and never before, derived from nothing about you or your
machine. It exists so the corpus can tell two contributions apart from one client
retrying. Delete it and you are a new contributor; nothing else changes.

## What it is not

Not a player, not a tagger, not a library manager, not a downloader. It does the
two things a CLAP embedding makes uniquely easy and stops.

## Why it exists

It is the **reference client** — the place the commons's contract is exercised end
to end — and not the way most people are expected to arrive. The argument for it
is in [`ADR-0009`](../../docs/decisions/ADR-0009-the-tool-is-useful-before-the-corpus-is.md):
a donation client with no local value has no first contributor, and this project
has measured proof that passive accumulation does not happen. What the tool does
locally is the draw; contributing to the [commons](https://clapback.seethroughlab.com)
is a byproduct of it.

The route to the commons for most people is the tool they already run.
[`ADR-0011`](../../docs/decisions/ADR-0011-the-commons-is-what-other-tools-plug-into.md)
says so: if you use beets, [`beets-clapback`](https://pypi.org/project/beets-clapback/)
does everything above against your library and can name what it contributes with
`mb_trackid`. If you are writing a tool, the contract this CLI follows is published
on its own as [`clapback-client`](https://pypi.org/project/clapback-client/) — no
dependency beyond the standard library, so a tool with its own embedder can take
part without ONNX Runtime. This CLI is built on it, which is what keeps the reference
client and the published contract from drifting apart.
