Metadata-Version: 2.3
Name: scigma_harvesting
Version: 0.3.0
Summary: Add your description here
Requires-Dist: pymarc>=5.3.1
Requires-Dist: python-decouple-typed>=3.11.0
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: pyzotero>=1.11.0
Requires-Dist: requests>=2.32.5
Requires-Python: >=3.11, <3.14
Description-Content-Type: text/markdown

# Harvesting for [SCIGMA](https://www.scigma.de/)

![License: AGPL-3.0](https://img.shields.io/badge/License-AGPL--3.0-blue.svg)

## Init

### Configuration

1. Install dependencies, for example with `uv sync`.
2. Copy `.env.example` to `.env`.

```bash
cp .env.example .env
```

3. Fill in your local API keys and IDs.

Initialize and run all production tests with:

```bash
uv run pytest -m prod -v
```

> [!NOTE]
> Running production tests (`-m prod`) will automatically download and initialize the STCV SQLite database if it does not exist or if a new version is available (checked via ETag).

4. Optional you may adapt `config.yaml` to your needs. It contains all configuration values.

## Usage

For working examples, see:

- [Minimal workflow](docs/examples/minimal_workflow.py) – Basic harvest pipeline with Zotero/JSON/RIS export.
- [Advanced raw queries](docs/examples/advanced_workflow.py) – Using raw query parameters (`base_raw`, `ddb_raw`, `stcv_raw`) for source-specific control.

### Top-Level Functions

The `harvest` package exposes the following functions:

- `harvest(...)` – Main search function. Queries all sources (STCV, BASE, DDB), deduplicates results, returns `list[HarvestRecord]`.
- `harvest_all(...)` – Like `harvest`, but without deduplication.
- `deduplicate(records, custom_filter=None)` – Deduplicate `HarvestRecord`s directly. Accepts optional `custom_filter` function for additional post-processing.
- `to_zotero(records, ...)` – Exports `HarvestRecord`s to Zotero (via API).
- `to_ris(records, path)` – Exports `HarvestRecord`s to a RIS file.
- `to_json(records, path)` – Exports `HarvestRecord`s to a JSON file.
- `activate_logging(debug=False)` – Configure logging: INFO level by default, DEBUG level when `debug=True`.

Source-specific search functions:

- `base_search(query, ...)` – Search BASE API directly.
- `ddb_search(query, ...)` – Search Deutsche Digitale Bibliothek directly.
- `stcv_search(must_terms, ...)` – Search STCV database directly.

### Data Model

The normalized data model `HarvestRecord` of the output is importable from the package and documented in [src/harvest/pipeline/models.py](src/harvest/pipeline/models.py).

### Source-Specific Functions and Raw Dumps

Each search function (`base_search`, `ddb_search`, `stcv_search`) accepts an optional **`debug_dump_path`** parameter to save raw API/database responses for debugging and reproducibility. This is especially useful for investigating query issues or preserving data snapshots.

Additionally, `stcv_search` supports a **`sql_query`** parameter for direct SQL queries against the STCV SQLite database. When provided, it bypasses all other filter parameters and executes the raw SQL, which must return rows containing a `cloi` column.

#### BASE (`scigma_harvesting.base`)

```python
from scigma_harvesting.base import search as base_search

# Single page search
results = base_search(
    query="dccreator: lossau",
    debug_dump_path=Path("dumps/base_response.json")
)
```

**Dump contents:** JSON file containing `source`, `query`, `params`, and `raw_xml_file` (path to separate `.xml` file with pretty-printed raw response)

**Special behavior:** Creates TWO files:

- `base_response.json` – Metadata and parsed results
- `base_response.xml` – Raw XML response from BASE API (pretty-printed)

#### DDB (`scigma_harvesting.ddb`)

```python
from scigma_harvesting.ddb import search as ddb_search

results = ddb_search(
    query="lossau",
    debug_dump_path=Path("dumps/ddb_response.json")
)
```

**Dump contents:** JSON file containing `source`, `query`, `params`, and `raw_response` (complete DDB API JSON response)

#### STCV (`scigma_harvesting.stcv`)

```python
from scigma_harvesting.stcv import search as stcv_search

# Standard search with filters
results = stcv_search(
    must_terms=["Aristoteles"],
    debug_dump_path=Path("dumps/stcv_response.json")
)

# Direct SQL query (must return rows with cloi column)
results = stcv_search(
    sql_query="SELECT cloi FROM title WHERE title_ti LIKE '%Aristoteles%'"
)
```

**Dump contents:** JSON file containing `source`, search parameters (`standalone`, `must_terms`, `must_not_terms`, `year_from`, `year_to`, `sql_query`), and `results` (parsed records)

**Note:** STCV uses a local SQLite database, so dumps contain the already-parsed results rather than raw API responses.

### Logging

The package uses Python's `logging` module. All modules (`base`, `ddb`, `stcv`, `pipeline`) log at **DEBUG** and **INFO** levels:

> [!WARNING]
> When `debug=True` HTTP request URLs and headers are logged, which may **include API keys**. Never use `debug=True` in production or in environments where logs are shared or persisted.

- **DEBUG**: Raw API requests/responses (query parameters, full XML/JSON responses), individual record details, Zotero batch creation summaries.
- **INFO**: High-level progress (search start/end, result counts, harvest completion, Zotero chunk processing).

To enable detailed logging in your own code, use the provided helper:

```python
from scigma_harvesting import activate_logging

# Activate logging (INFO level by default)
activate_logging()

# For DEBUG level (very verbose, including API requests/responses):
activate_logging(debug=True)
```

### Configuration Constants

The package uses several magic numbers for API limits and batch sizes. These can be modified in their respective modules:

| Constant                      | Module                              | Description                                                                                      | Default |
| ----------------------------- | ----------------------------------- | ------------------------------------------------------------------------------------------------ | ------- |
| `RATE_LIMIT`                  | `scigma_harvesting.config`          | Seconds between API requests (rate limiting)                                                     | `2`     |
| `MAX_HITS`                    | `scigma_harvesting.base.config`     | BASE API: max results per request (official limit)                                               | `120`   |
| `MAX_OFFSET`                  | `scigma_harvesting.base.config`     | BASE API: max offset value (official limit; combined with MAX_HITS, caps total at ~1080 results) | `999`   |
| `ZOTERO_CHUNK_SIZE`           | `scigma_harvesting.pipeline.config` | Zotero API: items per batch request (max allowed by Zotero)                                      | `50`    |
| `ZOTERO_CHUNK_DELAY`          | `scigma_harvesting.pipeline.config` | Delay between Zotero batch requests (seconds) to prevent rate limiting/timeouts                  | `2.0`   |
| `ZOTERO_HTTP_READ_TIMEOUT`    | `scigma_harvesting.pipeline.config` | HTTP read timeout for Zotero API requests (seconds)                                              | `60.0`  |
| `ZOTERO_HTTP_CONNECT_TIMEOUT` | `scigma_harvesting.pipeline.config` | HTTP connect timeout for Zotero API requests (seconds)                                           | `30.0`  |

The BASE API enforces a strict _1 query per second limit per API key_. Exceeding this will result in being blacklisted. The library's default `RATE_LIMIT=2` ensures compliance by adding a 2-second delay between requests. Do not reduce this value below 1 second.

Zotero API timeouts can be increased if you encounter frequent read or connection timeouts when exporting to Zotero.

## Troubleshooting

**Common issues:**

- **BASE API returns no results:** Verify `API_KEY_BASE` in `.env`. Check query syntax (Lucene/SOLR). BASE has a hard limit of ~1080 results per query.
- **Zotero export fails:** Verify `API_KEY_ZOTERO` and `USER_ID_ZOTERO`. Check collection key exists.
- **DDB IIIF/PDF links not accessible:** Many DDB resources require institutional authentication.

**Debug mode:**
Enable detailed logging or run development tests for debugging:

```python
from scigma_harvesting import activate_logging
activate_logging(debug=True)  # Enable DEBUG logging for all HTTP requests/responses
```

Or run dev tests to inspect individual components:

```bash
uv run pytest -m dev -v
```

**Debug logging in integration tests:**
To enable DEBUG logging for integration tests (e.g., to inspect HTTP requests), use the `--log-debug` flag:

```bash
# Without debug (default: no secrets in logs)
pytest tests/integration/

# With debug (WARNING: may log API keys and other secrets!)
pytest --log-debug tests/integration/
```

> **Note:** The `--log-debug` flag is disabled by default. When enabled, API keys may appear in log files (`tests/artifacts/logs/test_e2e_debug.log`). Never use `--log-debug` in CI or shared environments.

## Project Structure

```bash
src/harvest/
├── base/       # BASE API integration
├── ddb/        # Deutsche Digitale Bibliothek
├── stcv/       # Short Title Catalogue Vlaanderen
└── pipeline/   # Core harvest pipeline
    ├── core/   # Main pipeline module (harvest, harvest_all)
    ├── models  # HarvestRecord and Zotero export models
    └── export  # Export functions (to_zotero, to_ris, to_json)
```

## IO Quality

The `tests/integration/test_io_quality.py` module contains integration tests that verify data integrity during I/O operations (searching, processing, saving, loading). For manual comparision dumps can be found in `tests/artifacts/io-quality` after running this test.

### DDB

- API Access: The DDB search API is public and does not require an API key.
- MODS (MARC21) URLs: Generated via `/items/{id}/source/record` — publicly accessible.
- IIIF endpoints: Some IIIF manifests linked in DDB records may require institutional authentication tokens (not provided by this library).
- METS endpoints: Typically require authentication and are not publicly accessible.

#### Updating the Cache

The DDB uses a custom day-based timestamp system for dates. The `DDB_TIMEPARSER_CACHE` dictionary in `scigma_harvesting.ddb.ddb_cache` contains precomputed day timestamps for January 1st of each year.

If you need to update the DDB time parser cache (e.g., to extend it to cover more years):

1. Place your updated `ddb_timeparser_cache.tsv` file in `data/internal/ddb/`
2. Run the conversion script:

```bash
uv run python scripts/convert_ddb_cache.py data/internal/ddb/ddb_timeparser_cache.tsv src/scigma_harvesting/ddb/ddb_cache.py
```

This will generate a new `ddb_cache.py` file with the updated cache as a Python dictionary.

The TSV file format is simple: each line contains `<year>\t<ddb_day_timestamp>` where `<ddb_day_timestamp` is the time_stamp for the first day of the year.

## Development

### Git and Versioning

**GitFlow**

Merge non-squashing from `dev` to `main`. Fast-forwarding is not allowed in order to see the different releases on main as merge commit.

**Versioning**

Versioning is handled primarily via the version in `pyproject.toml`. To bump run

```bash
uv version --bump [beta|patch|minor|major]
```

Then commit and push.

In order to make a release from that, a `git tag` will trigger the CI:

```bash
# git tag
git tag -a $(uv version | awk '{print $2}') -m "Release $(uv version | awk '{print $2}')"
# git push tag
git push origin $(uv version | awk '{print $2}')
```

### Styling

The code is formated by [black](https://pypi.org/project/black/) formatter.

### Tests

Run all tests (including dev tests, not recommended for production!) with:

```bash
uv run pytest -v
```

**Test markers:**

- `-m prod` – Production tests (safe for CI/CD, includes STCV database initialization).
- `-m dev` – Development/debug tests only (not suitable for production).

If errors occur, you can run individual tests or test modules for debugging:

```bash
# Run a specific test file
uv run pytest tests/unit/pipeline/test_pipeline.py -v

# Run a single test function
uv run pytest tests/unit/pipeline/test_pipeline.py::test_base_raw_passed_through_directly -v
```
