Metadata-Version: 2.5
Name: graflo
Version: 1.16.0
Summary: Manifest-driven Graph Schema & Transformation Language (GSTL): declare a labeled property graph with explicit identities, typed properties and semantic grounding; evolve it with an invertible op algebra, content-addressed commits and three-way merge; ingest from CSV/JSON/Parquet/SQL/RDF/SPARQL/API/Kafka; and project to ArangoDB, Neo4j, TigerGraph, FalkorDB, Memgraph, NebulaGraph, PostgreSQL or a chunked file backend.
Project-URL: Changelog, https://github.com/growgraph/graflo/blob/main/CHANGELOG.md
Project-URL: Documentation, https://growgraph.github.io/graflo
Project-URL: Homepage, https://github.com/growgraph/graflo
Author-email: Alexander Belikov <alexander@growgraph.dev>
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Requires-Python: >=3.11
Requires-Dist: boto3<2,>=1.35.0
Requires-Dist: click<9,>=8.2.0
Requires-Dist: confluent-kafka>=2.6.0
Requires-Dist: falkordb>=1.0.9
Requires-Dist: ijson<4,>=3.2.3
Requires-Dist: nebula3-python>=3.8.3
Requires-Dist: nebula5-python>=5.2.1
Requires-Dist: neo4j<6,>=5.22.0
Requires-Dist: networkx~=3.3
Requires-Dist: pandas-stubs==2.3.0.250703
Requires-Dist: pandas<4,>=2.0.3
Requires-Dist: psycopg2-binary>=2.9.11
Requires-Dist: pydantic-settings>=2.12.0
Requires-Dist: pydantic>=2.12.5
Requires-Dist: pymgclient>=1.3.1
Requires-Dist: python-arango<9,>=8.1.2
Requires-Dist: rdflib>=7.0.0
Requires-Dist: redis>=5.0.0
Requires-Dist: requests>=2.31.0
Requires-Dist: sparqlwrapper>=2.0.0
Requires-Dist: sqlalchemy>=2.0.0
Requires-Dist: strenum>=0.4.15
Requires-Dist: suthing<0.7,>=0.6.1
Requires-Dist: urllib3>=2.0.0
Requires-Dist: xmltodict<0.15,>=0.14.2
Provides-Extra: dev
Requires-Dist: hypothesis>=6.140.0; extra == 'dev'
Requires-Dist: pre-commit>=4.2.0; extra == 'dev'
Requires-Dist: pytest-timeout>=2.4.0; extra == 'dev'
Requires-Dist: pytest-xdist>=3.8.0; extra == 'dev'
Requires-Dist: pytest>=9.0.2; extra == 'dev'
Requires-Dist: ty>=0.0.1a24; extra == 'dev'
Provides-Extra: docs
Requires-Dist: mkdocs-gen-files>=0.5.0; extra == 'docs'
Requires-Dist: mkdocs-glightbox>=0.4.0; extra == 'docs'
Requires-Dist: mkdocs-literate-nav>=0.6.2; extra == 'docs'
Requires-Dist: mkdocs-material>=9.6.12; extra == 'docs'
Requires-Dist: mkdocs-table-reader-plugin>=3.1.0; extra == 'docs'
Requires-Dist: mkdocstrings[python]>=0.29.1; extra == 'docs'
Requires-Dist: properdocs>=1.6.7; extra == 'docs'
Provides-Extra: plot
Requires-Dist: pygraphviz<3,>=2.0; extra == 'plot'
Description-Content-Type: text/markdown

# GraFlo <img src="https://raw.githubusercontent.com/growgraph/graflo/main/docs/assets/favicon.ico" alt="graflo logo" style="height: 32px; width:32px;"/>

![Python](https://img.shields.io/badge/python-3.11%2B-blue.svg)
[![PyPI version](https://badge.fury.io/py/graflo.svg)](https://badge.fury.io/py/graflo)
[![PyPI Downloads](https://static.pepy.tech/badge/graflo)](https://pepy.tech/projects/graflo)
[![Docs](https://img.shields.io/badge/docs-growgraph.github.io-orange.svg)](https://growgraph.github.io/graflo)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-green)](https://github.com/growgraph/graflo/blob/main/LICENSE)
[![pre-commit](https://github.com/growgraph/graflo/actions/workflows/pre-commit.yml/badge.svg)](https://github.com/growgraph/graflo/actions/workflows/pre-commit.yml)
[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.15446131.svg)](https://doi.org/10.5281/zenodo.15446131)

GraFlo is a Python library that turns records from files, SQL databases, RDF,
REST APIs or Kafka topics into a labeled property graph. You describe the graph
once, in a YAML file called a manifest, and GraFlo creates the schema and
writes the vertices and edges into the graph database of your choice, or into
a directory on disk.

It is for engineers who build a graph from several sources and want its
description in one reviewable file rather than spread across load scripts.

## What you can do with it

- **Describe a graph once and load data into it.** A manifest names the vertex
  and edge types, says which properties identify a vertex, and says how each
  kind of record becomes vertices and edges. The same manifest loads into
  ArangoDB, Neo4j, TigerGraph, FalkorDB, Memgraph, NebulaGraph, PostgreSQL or
  the file backend, and records with the same identity become one vertex.
  GraFlo also copies an existing graph from Neo4j, ArangoDB or PostgreSQL into
  another database (`GraphEngine.migrate_graph`).
- **Change the description over time, with a recorded history.** Renaming a
  type, combining two types or changing a property type is a typed operation.
  Operations are recorded as commits (`graflo commit`, `log`, `checkout`,
  `verify`, `revert`) that you can replay, check and, for most operations,
  undo. Two
  branches of changes to one manifest are reconciled with a three-way merge
  (`graflo merge3`), and two manifests written by different teams are combined
  into one with a union (`graflo merge`).
- **Check and infer descriptions.** GraFlo infers a manifest from a PostgreSQL
  database or an OWL ontology, proposes the properties that identify a record
  from sample data, and checks a manifest against a conformance profile
  (`graflo check`), a set of modeling rules such as "every vertex type
  declares its identity".

## A taste

A manifest has three blocks: `schema` says what the graph looks like,
`ingestion_model` says how records map onto it, and `bindings` says where the
records come from. This one reads CSV files with the columns `person_id`,
`person` and `department`:

```yaml
schema:
    metadata: {name: hr}
    graph:
        vertex_config:
            vertices:
            -   {name: person, properties: [id, name], identity: [id]}
            -   {name: department, properties: [name], identity: [name]}
        edge_config:
            edges: [{source: person, target: department}]
ingestion_model:
    resources:
    -   name: departments
        pipeline:
        -   {vertex: person, from: {id: person_id, name: person}}
        -   {vertex: department, from: {name: department}}
bindings:
    connectors:
    -   {regex: "^dep.*\\.csv$", sub_path: data, resource_name: departments}
```

This loads it into ArangoDB:

```python
from graflo import GraphEngine, GraphManifest
from graflo.connections import ArangoConfig

manifest = GraphManifest.from_yaml("manifest.yaml")
manifest.finish_init()
engine = GraphEngine()
engine.define_and_ingest(manifest=manifest, target_db_config=ArangoConfig.from_env())
```

`ArangoConfig.from_env()` reads `ARANGO_URI`, `ARANGO_USERNAME`,
`ARANGO_PASSWORD` and `ARANGO_DATABASE`; every database has such a class. See
[Database connections](https://growgraph.github.io/graflo/guides/database_connections/).

## Documentation

Full documentation: [growgraph.github.io/graflo](https://growgraph.github.io/graflo)

- [Quick start](https://growgraph.github.io/graflo/getting_started/quickstart/): two CSV files into a graph, step by step
- [Creating a manifest](https://growgraph.github.io/graflo/getting_started/creating_manifest/): the three blocks of a manifest
- [Examples](https://growgraph.github.io/graflo/examples/): runnable examples, one question each, with their data under [`examples/`](https://github.com/growgraph/graflo/tree/main/examples)
- [Concepts](https://growgraph.github.io/graflo/concepts/): schema, identity, ingestion, connectors, evolution and version control
- [Guides](https://growgraph.github.io/graflo/guides/): database connections, graph migration, schema inference, API wiring, bulk load
- [GraFlo ontology](https://growgraph.github.io/graflo/concepts/schema/ontology/): a manifest as RDF (`graflo manifest-to-rdf`, `graflo rdf-to-manifest`)

## Installation

GraFlo needs Python 3.11 or newer. The database clients, RDF and Kafka support
are part of the default install.

```bash
pip install graflo
```

Optional extras (see the
[Installation](https://growgraph.github.io/graflo/getting_started/installation/) guide):

- `dev`: pytest and its plugins, hypothesis, ty, pre-commit
- `docs`: ProperDocs and its plugins, for building the documentation site
- `plot`: `pygraphviz` for `graflo plot-manifest` and the `--plot` figures of
  `graflo merge` and `graflo merge3`

```bash
pip install "graflo[dev,docs,plot]"
```

## Development

To install from a clone:

```shell
git clone git@github.com:growgraph/graflo.git && cd graflo
uv sync --extra dev
```

See the [Contributing Guide](https://growgraph.github.io/graflo/contributing/) for the full workflow.

### Tests

The database tests need the database containers. Start them from a clone with
the scripts under [docker/](https://github.com/growgraph/graflo/tree/main/docker):

```shell
cd docker
./start-all.sh    # Start all services
./stop-all.sh     # Stop all services
./cleanup-all.sh  # Remove containers and volumes
```

Per-engine compose files and ports are documented in the
[docker README](https://github.com/growgraph/graflo/blob/main/docker/README.md).

To run the tests:

```shell
uv run pytest test
```

TigerGraph, NebulaGraph and Kafka tests are skipped unless you pass
`--run-tigergraph`, `--run-nebula` or `--run-kafka`.

The suites that need no database run without the containers, and CI runs them
on every pull request:

```shell
uv run pytest test --ignore=test/db --ignore=test/data_source --ignore=test/object_storage
```

## License

Open source under the [Apache License 2.0](https://github.com/growgraph/graflo/blob/main/LICENSE).
Copyright and trademark notices are in
[NOTICE](https://github.com/growgraph/graflo/blob/main/NOTICE): the license grants no rights in the
**GraFlo** and **GrowGraph** marks. Releases before the relicensing shipped under the Business
Source License 1.1 and keep those terms; see the
[changelog](https://github.com/growgraph/graflo/blob/main/CHANGELOG.md).

## Contributing

Contributions are welcome. See the
[Contributing Guide](https://growgraph.github.io/graflo/contributing/). Contributors accept the
[Contributor License Agreement](https://github.com/growgraph/graflo/blob/main/CLA.md) once, by
commenting on their first pull request.
