Metadata-Version: 2.5
Name: bluecore-models
Version: 0.34.0
Summary: Blue Core BIBFRAME Data Models
License-File: LICENSE
Requires-Python: >=3.12
Requires-Dist: alembic>=1.14.1
Requires-Dist: psycopg[binary]>=3.2
Requires-Dist: pyld>=2.0.4
Requires-Dist: rdflib>=7.1.3
Requires-Dist: sqlalchemy[asyncio]>=2.1.1
Requires-Dist: tenacity>=9.1.4
Description-Content-Type: text/markdown

# Blue Core Data Models
The Blue Core Data Models are used in [Blue Core API](https://github.com/blue-core-lod/bluecore_api) 
and in the [Blue Core Workflows](https://github.com/blue-core-lod/bluecore-workflows) services.  

## 🐳 Run Postgres with Docker
To run the Postgres with the Blue Core Database, run the following command from this directory:

`docker run --name bluecore_db -e POSTGRES_USER=airflow -e POSTGRES_PASSWORD=airflow -v ./create-db.sql:/docker-entrypoint-initdb.d/create_database.sql -p 5432:5432 postgres:17`

---

## 🛠️ Installing
- Install via pip: `pip install bluecore-models`
- Install via uv: `uv add bluecore-models`

---

## 🗄️ Database Management
The [SQLAlchemy](https://www.sqlalchemy.org/) Object Relational Mapper (ORM) is used to create
the Bluecore database models. 

```mermaid
erDiagram
    ResourceBase ||--o{ Hub : "has"
    ResourceBase ||--o{ Instance : "has"
    ResourceBase ||--o{ Work : "has"
    ResourceBase ||--o{ OtherResource : "has"
    ResourceBase ||--o{ Profile : "has"
    ResourceBase ||--o{ ResourceBibframeClass : "has classes"
    ResourceBase ||--o{ Version : "has versions"
    ResourceBase ||--o{ BibframeOtherResources : "has other resources"

    Hub ||--o{ Work : "has"
    Work ||--o{ Instance : "has"
    
    BibframeClass ||--o{ ResourceBibframeClass : "classifies"
    
    OtherResource ||--o{ BibframeOtherResources : "links to"

    Profile ||--o{ ProfileRelation : "nests"
    Profile ||--o{ ProfileRelation : "is nested by"
```

### Profile nesting

A Profile is a Sinopia profile: JSON-LD describing how to edit a kind of resource. Its
`data` is stored unframed, because Sinopia Editor needs back the shape it sent.

A profile's data may name other profiles that it nests, with
`sinopia:hasResourceTemplateId`. Those references are stored as profile URIs, and each
one that resolves is recorded as a `ProfileRelation` row with a foreign key on both
ends. The nesting is many-to-many — one profile is commonly nested by several others —
so it cannot be a column on `profiles`.

Both foreign keys cascade, so deleting either end removes the edge, and a `CHECK`
constraint stops a profile nesting itself. That means "is this profile nested by
anything?" is derived rather than stored, and cannot go stale:

```python
# Every top-level profile, as one SQL statement with an inlined NOT EXISTS.
session.scalars(select(Profile).where(~Profile.is_nested))
```

`Profile.is_nested` is deferred, so an ordinary `select(Profile)` does not carry the
subquery. It is meant for filtering in SQL: reading it off an instance that did not
select it issues a query, so it is an N+1 in a loop. Use `Profile.children` and
`Profile.parents` to walk the relation itself.

References are expected to be profile URIs. A reference naming nothing stored records
no row rather than raising — bluecore_api rejects those on save, where there is a
cataloger to tell.

Works are linked to Instances by `bf:instanceOf` / `bf:hasInstance`, and to a Hub by
`bf:expressionOf` (see `BluecoreGraph._link`).

Both ends of a link need a URI. A resource created in an editor arrives without one
and is minted a Bluecore URI, but a blank node in a bulk-loaded record is an inline
description of something the record merely refers to — LC catalog data often states
`bf:expressionOf` against an anonymous `bf:Hub` — so it gets no record and no link
(see `BluecoreGraph._anonymous_description`).

### Database Migrations with Alembic
The [Alembic](https://alembic.sqlalchemy.org/en/latest/) database migration package is used
to manage database changes with the Bluecore Data models.

To create a new migration, ensure that the Postgres database is available and then run:
- `uv run alembic revision --autogenerate -m "{short message describing change}`

A new migration script will be created in the `bluecore_store_migration` directory. Be sure
to add the new script to the repository with `git`.

#### Applying Migrations
To apply all of the migrations, run the following command:
- `uv run alembic upgrade head`

---

## 🧹 Linter for Python 
bluecore-models uses [ruff](https://docs.astral.sh/ruff/)
- `uv run ruff check`

To auto-fix errors in both (where possible):
- `uv run ruff check --fix`

Check formatting differences without changing files:
- `uv run ruff format --diff`

Apply Ruff's code formatting:
- `uv run ruff format`

---

## 🧪 Running Tests
The test suite is written using pytest and is executed via uv.
All tests are located in the tests/ directory.

#### Run All Tests
`uv run pytest`

#### Run a specific test file
`uv run pytest tests/test_models.py`

#### Run a specific test function
`uv run pytest tests/test_models.py -k test_updated_instance`

#### Show output (prints/logs) during test execution
`uv run pytest -s`

💡 Make sure your virtual environment is activated and dependencies are installed with uv before running tests.

---

## 📊 Benchmarking `save_graph`

`benchmarks/save_graph_bench.py` persists a set of Bibframe graphs through the
real save path (URI minting, resource save, linking, bf-class updates) and
reports throughput (graphs/s, triples/s). With `--profile` it prints a cProfile
hot-spot report, which is handy for finding where the save path spends its time.

It writes to a Postgres, so first start one — the [Run Postgres with Docker](#-run-postgres-with-docker)
command above works (it creates a `bluecore` database). Then point the benchmark
at it with `--database-url` (or the `DATABASE_URL` env var); the benchmark
creates the schema if it isn't there.

#### Run the benchmark (50 saves of the sample graphs)
```
uv run python benchmarks/save_graph_bench.py \
  --database-url postgresql+psycopg://airflow:airflow@localhost:5432/bluecore \
  --count 50
```

#### Print a cProfile hot-spot report
Add `--profile`:
```
uv run python benchmarks/save_graph_bench.py \
  --database-url postgresql+psycopg://airflow:airflow@localhost:5432/bluecore \
  --count 50 --profile
```

Other options:
- `--reset` — `TRUNCATE` the resource tables first (don't point this at data you care about).
- `--input "<glob>"` — RDF files to load (defaults to `tests/data/*.jsonld`); pass a
  directory of `.rdf`/`.jsonld` records for a larger, more representative run.

💡 Since the graph content doesn't change *which* code runs, reusing a few sample
records many times (`--count`) is a fair, repeatable way to measure changes to
the save path (e.g. before/after an optimization).

---

## ⬆️ Publishing to Pypi
To publish the `bluecore-models` to [pypi](https://pypi.org/project/bluecore-models/), the
following steps need to be taken. 

1. Update the version in `pyproject.toml` either in a feature branch PR or in a
   dedicated PR.
2. After the PR is merged, create a [tagged release](https://github.com/blue-core-lod/bluecore-models/releases) 
   using the same version (prepended with a `v` i.e. `v0.4.2`.
3. Once the tagged release is saved, the [Publish to PyPi](https://github.com/blue-core-lod/bluecore-models/actions/workflows/publish.yml)
   Github Action should then publish the release to PyPi. 
