Metadata-Version: 2.4
Name: goldrush
Version: 0.4.1
Summary: Generate Gold Rush (CoAlliance) matchKey for pymarc or mrrc records
Author: Ed Summers
Author-email: Ed Summers <ehs@pobox.com>
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Requires-Dist: mrrc>=0.9.1 ; extra == 'mrrc'
Requires-Dist: pymarc>=5.3.1 ; extra == 'pymarc'
Requires-Python: >=3.10
Provides-Extra: mrrc
Provides-Extra: pymarc
Description-Content-Type: text/markdown

# goldrush

**Note: This is alpha software. If you rely on this library you should do so with the
understanding that you might find errors in the Gold Rush key that is generated.
Lets make it better together!**

---

[![Build status](https://gitlab.com/pymarc/goldrush/badges/main/pipeline.svg)](https://gitlab.com/pymarc/goldrush/-/commits/main)

*goldrush* is a Python implementation of the [Gold Rush match key algorithm] for
identifying "duplicate" MARC records, or records that appear to be about the
same bibliographic item. It provides a function that you pass a MARC record and
get back the Gold Rush key as a string. Both [pymarc] and [mrrc] records are
supported — goldrush is duck-typed and imports neither backend itself, so you
use whichever one you already parse records with.

goldrush is a class-for-class port of the canonical, production-verified Java
reference implementation, [coalliance-matchkey], maintained by the Colorado
Alliance of Research Libraries (CoAlliance). It targets algorithm version
**`_v07312026`**, whose version string is embedded in every generated key (just
before the trailing format character) so you can tell which algorithm produced a
given key. See [`docs/CoAlliance_Match_Key.md`](docs/CoAlliance_Match_Key.md)
for the field-by-field specification.

The initial code was found in [pymarc_dedupe] created by Max Kadel and Jane
Sandberg at Princeton University Library. The goldrush module was created
because pymarc_dedupe included other functionality that was unrelated to Gold
Rush, and pymarc_dedupe was not installable as a module via PyPI.

## Usage

First install goldrush together with a MARC backend. goldrush does not depend on
either [pymarc] or [mrrc] directly, so pick one via an extra:

```
pip install goldrush[pymarc]     # or: pip install goldrush[mrrc]
```

Then load a record with your backend and generate a key:

```python
>>> from goldrush import goldrush
>>> from pymarc import MARCReader          # or: from mrrc import MARCReader
>>> record = next(MARCReader(open('marc.dat', 'rb')))
>>> goldrush(record)
'pragmaticprogrammerfromjourneymantomaster___________________________________________________________2000____1__addisa________________________________________hunt_________________v07312026p'
```

The same call works with an [mrrc] record — its differential test in
`tests/test_corpus.py` runs the whole corpus through both backends and confirms
byte-identical keys.

The CoAlliance indexer optionally uses the MARC *filename* as a hint when
deciding whether a record is electronic or print. To match that behaviour, pass
the filename:

```python
>>> goldrush(record, marc_filename_hint='cu-electronic-2026-04.marc')
```

Pass nothing (the default) for pure-MARC format detection.

## Matching records

Two records match when their keys are equal — comparing keys is all the matching
there is. [`docs/examples/duplicates.mrc`](docs/examples/duplicates.mrc) holds
three real records for one 2023 CRC Press title: the print edition as cataloged
by the Library of Congress, the same printed book as cataloged independently in
K10plus, and the e-book as cataloged by the Library of Congress.

```python
>>> from goldrush import goldrush
>>> from pymarc import MARCReader
>>> with open('docs/examples/duplicates.mrc', 'rb') as fh:
...     lc_print, k10_print, lc_electronic = [goldrush(r) for r in MARCReader(fh)]
>>> lc_print == k10_print
True
>>> lc_print == lc_electronic
False
```

The two print records match even though the catalogers disagreed about the
leading article (`$aA pragmatic programmer` vs `$aThe pragmatic programmer`),
whether to record an edition statement, whether to bracket the date (`[2023]` vs
`2023`), and how to punctuate the publisher — Gold Rush normalizes all of that
away.

The e-book does not match, and the reason is visible in the keys: they differ at
exactly one character, the trailing format byte.

```python
>>> lc_print[-10:]
'v07312026p'
>>> lc_electronic[-10:]
'v07312026e'
>>> lc_print[:-1] == lc_electronic[:-1]
True
```

See [`docs/examples/README.md`](docs/examples/README.md) for record provenance;
`tests/test_examples.py` asserts these relationships under both backends.

### Masking parts of the key

Sometimes a component of the key encodes a distinction you do not care about for
the task in front of you. Because the key is a fixed-width positional string —
every component occupies the same character range in every key — you can ignore a
component by slicing it out or blanking it before you compare. Nothing in
goldrush does this for you; it is your string, and masking is ordinary Python.

The print/electronic distinction is a good worked example. Keeping print and
electronic apart is usually what you want, but if you are clustering at the work
level — "how many editions of this do we hold, in any carrier?" — you can mask it
out. The format character is the last character of the key, so dropping it is the
whole recipe:

```python
>>> def format_agnostic(key):
...     """Gold Rush key with the trailing print/electronic byte removed."""
...     return key[:-1]
...
>>> format_agnostic(lc_print) == format_agnostic(lc_electronic)
True
>>> len({format_agnostic(k) for k in (lc_print, k10_print, lc_electronic)})
1
```

All three example records collapse to a single masked key.

The same move works on any other component, depending on what you are doing.
These are the ranges, as generated by `generator.generate`:

| Component | Range | Width |
| --- | --- | --- |
| title | `0:95` | 95 |
| GMD (disabled since 2022, always underscores) | `95:100` | 5 |
| publication year | `100:104` | 4 |
| pagination | `104:108` | 4 |
| edition | `108:111` | 3 |
| publisher | `111:116` | 5 |
| leader type | `116:117` | 1 |
| title part | `117:147` | 30 |
| title number | `147:157` | 10 |
| author | `157:162` | 5 |
| title dates | `162:177` | 15 |
| algorithm version | `177:187` | 10 |
| format (`p`/`e`) | `187:188` | 1 |

So: blank `104:108` if you want records to match across differing pagination —
common between a print record and an `$a1 online resource` record, and in fact
the format byte alone is often not enough to collapse print and electronic for
that reason. Blank `100:104` to ignore a one-year publication-date disagreement.
Blank `157:162` if you are matching on title and imprint alone. Keep the width
when you blank, so later components stay aligned:

```python
>>> def work_level(key):
...     """Key with both pagination and the format byte masked out."""
...     return key[:104] + '____' + key[108:-1]
...
>>> len({work_level(k) for k in (lc_print, k10_print, lc_electronic)})
1
```

Masking widens matches, so it also admits false positives — the more you blank,
the more genuinely different manifestations collapse together. Note also that a
masked key is no longer a Gold Rush key, so don't store one where something else
expects the real thing. See
[`docs/CoAlliance_Match_Key.md`](docs/CoAlliance_Match_Key.md) for what each
component is derived from.

## Verifying against the reference

The reference implementation is the behavioural contract. The tests port its
JUnit suite to pytest and additionally diff goldrush's output for every record in
`tests/marc.dat` against a golden file (`tests/marc.dat.keys`) generated by the
reference `coa_matchkey_v07312026.jar`. To regenerate the golden file (requires
Java and marc4j 2.9.6 on the classpath):

```
java -Dorg.coalliance.indexing.fileName= \
     -cp coa_matchkey_v07312026.jar:marc4j-2.9.6.jar \
     org.coalliance.matchkey.cli.MatchKeyCli tests/marc.dat > tests/marc.dat.keys
```

## Credits and licensing

goldrush is licensed under the [Apache License, Version 2.0](LICENSE).

The Gold Rush matchKey algorithm is the work of the
[Colorado Alliance of Research Libraries], and goldrush's field-extraction and
normalization logic (`src/goldrush/fields/`, `src/goldrush/util/`,
`generator.py`, and `version.py`) is a port of their Apache-2.0 reference
implementation, [coalliance-matchkey] (Copyright 2026 Colorado Alliance of
Research Libraries). The ISBN-validation logic additionally derives from the
solrmarc-marc4j project. These attributions are recorded in [NOTICE](NOTICE).

[Gold Rush match key algorithm]: https://coalliance.org/sites/default/files/GoldRush-Match_KeyJanuary2024_0.doc
[coalliance-matchkey]: https://github.com/co-alliance/coalliance-matchkey
[Colorado Alliance of Research Libraries]: https://coalliance.org
[pymarc_dedupe]: https://github.com/pulibrary/pymarc_dedupe
[pymarc]: https://gitlab.com/pymarc/pymarc
[mrrc]: https://github.com/dchud/mrrc
