Metadata-Version: 2.5
Name: ontocast
Version: 0.6.5
Summary: Agentic ontology and knowledge graph co-generation
Project-URL: Changelog, https://github.com/growgraph/ontocast/blob/main/CHANGELOG.md
Project-URL: Documentation, https://growgraph.github.io/ontocast
Project-URL: Homepage, https://github.com/growgraph/ontocast
Project-URL: Issues, https://github.com/growgraph/ontocast/issues
Project-URL: Repository, https://github.com/growgraph/ontocast
Author-email: Alexander Belikov <alexander@growgraph.dev>
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: agent,knowledge-graph,llm,ontology,rdf,semantic-web,sparql,triple-extraction
Classifier: Development Status :: 4 - Beta
Classifier: Framework :: FastAPI
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: <4.0,>=3.12
Requires-Dist: httpx<1,>=0.27.0
Requires-Dist: langchain-core<2,>=1.0
Requires-Dist: langgraph<2,>=1.0
Requires-Dist: numpy<3,>=1.26
Requires-Dist: oxrdflib<0.6,>=0.5.0
Requires-Dist: pydantic-settings<3,>=2.5.0
Requires-Dist: pydantic<3,>=2.11.9
Requires-Dist: pyld<4,>=3.0
Requires-Dist: pyoxigraph<0.6,>=0.5.8
Requires-Dist: pyyaml<7,>=6.0.2
Requires-Dist: rdflib<8,>=7.1.4
Requires-Dist: scikit-learn<2,>=1.4
Provides-Extra: all
Requires-Dist: click<9,>=8.1.8; extra == 'all'
Requires-Dist: docling-core[chunking]<3,>=2.57.0; extra == 'all'
Requires-Dist: docling>=2.57.0; extra == 'all'
Requires-Dist: duckduckgo-search>=8.1.1; extra == 'all'
Requires-Dist: easyocr>=1.7.2; extra == 'all'
Requires-Dist: fastapi<1,>=0.115.0; extra == 'all'
Requires-Dist: fastembed<1,>=0.8.0; extra == 'all'
Requires-Dist: hdbscan>=0.8.41; extra == 'all'
Requires-Dist: lancedb>=0.33.0; extra == 'all'
Requires-Dist: langchain-anthropic<2,>=1.0; extra == 'all'
Requires-Dist: langchain-google-genai<5,>=4.0; extra == 'all'
Requires-Dist: langchain-ollama<2,>=1.0; extra == 'all'
Requires-Dist: langchain-openai<2,>=1.0; extra == 'all'
Requires-Dist: networkx<4,>=3.0; extra == 'all'
Requires-Dist: pyshacl>=0.30.0; extra == 'all'
Requires-Dist: python-dotenv<2,>=1.0.0; extra == 'all'
Requires-Dist: python-multipart<1,>=0.0.22; extra == 'all'
Requires-Dist: qdrant-client<2,>=1.15.1; extra == 'all'
Requires-Dist: rich<16,>=14.0.0; extra == 'all'
Requires-Dist: sentence-transformers>=5.1.1; extra == 'all'
Requires-Dist: starlette<2,>=1.0; extra == 'all'
Requires-Dist: torch>=2.0; extra == 'all'
Requires-Dist: umap-learn>=0.5.12; extra == 'all'
Requires-Dist: uvicorn[standard]<1,>=0.32.0; extra == 'all'
Provides-Extra: anthropic
Requires-Dist: langchain-anthropic<2,>=1.0; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: anyio>=4.4.0; extra == 'dev'
Requires-Dist: pre-commit>=4.2.0; extra == 'dev'
Requires-Dist: pytest-order>=1.3.0; extra == 'dev'
Requires-Dist: pytest>=8.3.5; extra == 'dev'
Requires-Dist: ruff>=0.11.2; extra == 'dev'
Requires-Dist: ty>=0.0.14; extra == 'dev'
Provides-Extra: doc-processing
Requires-Dist: docling-core[chunking]<3,>=2.57.0; extra == 'doc-processing'
Requires-Dist: docling>=2.57.0; extra == 'doc-processing'
Requires-Dist: easyocr>=1.7.2; extra == 'doc-processing'
Requires-Dist: sentence-transformers>=5.1.1; extra == 'doc-processing'
Provides-Extra: docs
Requires-Dist: mkdocs-gen-files>=0.5.0; extra == 'docs'
Requires-Dist: mkdocs-glightbox>=0.4.0; extra == 'docs'
Requires-Dist: mkdocs-literate-nav>=0.6.2; extra == 'docs'
Requires-Dist: mkdocs-material>=9.6.14; extra == 'docs'
Requires-Dist: mkdocstrings-python>=2.0.3; extra == 'docs'
Requires-Dist: mkdocstrings[python]>=0.29.1; extra == 'docs'
Requires-Dist: properdocs>=1.6.7; extra == 'docs'
Provides-Extra: documents
Requires-Dist: docling-core[chunking]<3,>=2.57.0; extra == 'documents'
Provides-Extra: google
Requires-Dist: langchain-google-genai<5,>=4.0; extra == 'google'
Provides-Extra: graph
Requires-Dist: networkx<4,>=3.0; extra == 'graph'
Provides-Extra: lancedb
Requires-Dist: fastembed<1,>=0.8.0; extra == 'lancedb'
Requires-Dist: lancedb>=0.33.0; extra == 'lancedb'
Provides-Extra: ollama
Requires-Dist: langchain-ollama<2,>=1.0; extra == 'ollama'
Provides-Extra: openai
Requires-Dist: langchain-openai<2,>=1.0; extra == 'openai'
Provides-Extra: plot
Requires-Dist: pygraphviz<3,>=2.0; extra == 'plot'
Provides-Extra: qdrant
Requires-Dist: fastembed<1,>=0.8.0; extra == 'qdrant'
Requires-Dist: qdrant-client<2,>=1.15.1; extra == 'qdrant'
Provides-Extra: semantic-chunking
Requires-Dist: hdbscan>=0.8.41; extra == 'semantic-chunking'
Requires-Dist: sentence-transformers>=5.1.1; extra == 'semantic-chunking'
Requires-Dist: torch>=2.0; extra == 'semantic-chunking'
Requires-Dist: umap-learn>=0.5.12; extra == 'semantic-chunking'
Provides-Extra: server
Requires-Dist: click<9,>=8.1.8; extra == 'server'
Requires-Dist: docling-core[chunking]<3,>=2.57.0; extra == 'server'
Requires-Dist: fastapi<1,>=0.115.0; extra == 'server'
Requires-Dist: python-dotenv<2,>=1.0.0; extra == 'server'
Requires-Dist: python-multipart<1,>=0.0.22; extra == 'server'
Requires-Dist: rich<16,>=14.0.0; extra == 'server'
Requires-Dist: starlette<2,>=1.0; extra == 'server'
Requires-Dist: uvicorn[standard]<1,>=0.32.0; extra == 'server'
Provides-Extra: shacl
Requires-Dist: pyshacl>=0.30.0; extra == 'shacl'
Provides-Extra: sparse
Requires-Dist: fastembed<1,>=0.8.0; extra == 'sparse'
Provides-Extra: web-search
Requires-Dist: duckduckgo-search>=8.1.1; extra == 'web-search'
Description-Content-Type: text/markdown

# OntoCast <img src="https://raw.githubusercontent.com/growgraph/ontocast/refs/heads/main/docs/assets/logo.png" alt="OntoCast logo" style="height: 32px; width:32px;"/>

**Ontology-guided extraction of RDF knowledge graphs from documents.**

![Python](https://img.shields.io/badge/python-3.12%2B-blue.svg)
[![PyPI version](https://badge.fury.io/py/ontocast.svg)](https://badge.fury.io/py/ontocast)
[![PyPI Downloads](https://static.pepy.tech/badge/ontocast)](https://pepy.tech/projects/ontocast)
[![Docs](https://img.shields.io/badge/docs-growgraph.github.io-224777.svg)](https://growgraph.github.io/ontocast/)
[![License](https://img.shields.io/badge/License-Apache_2.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
[![pre-commit](https://github.com/growgraph/ontocast/actions/workflows/pre-commit.yml/badge.svg)](https://github.com/growgraph/ontocast/actions/workflows/pre-commit.yml)
[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.17796467.svg)](https://doi.org/10.5281/zenodo.17796467)

OntoCast reads documents and writes an RDF knowledge graph: an ontology that
describes the domain, and the facts the documents state in its terms. Give it
your ontologies and it extracts facts against them; give it none and it builds
one as it reads. Run it as an HTTP service, as a batch command, or inside your
own LangChain or LangGraph agent.

**Documentation:** [growgraph.github.io/ontocast](https://growgraph.github.io/ontocast/)

## Features

- **Ontology and facts together.** Each part of a document goes through a
  language model in a render-and-critique loop, in parallel; ontology changes
  are merged, versioned and checked before facts are extracted against them.
- **Patches, not rewrites.** The model emits insert/delete updates to a graph
  rather than regenerating it.
- **Entity disambiguation.** Mentions of the same entity across a document are
  merged into one.
- **Validation.** Deterministic checks, SHACL shapes, and repairs that need no
  extra model call.
- **Provenance.** Facts carry RDF 1.2 provenance back to the text, which you can
  strip on output.
- **Ontology context.** Each part of a document is shown the ontology it needs:
  chosen from your catalog, retrieved from a vector store (LanceDB or Qdrant),
  or fixed.
- **Storage.** In memory by default, Apache Jena Fuseki for persistence,
  partitioned by tenant and project.
- **A light core.** The base install embeds without a document-processing stack
  or an ML runtime; those are extras.

## Install

```sh
pip install "ontocast[server,openai,doc-processing]"
```

`server` provides the `ontocast` command and HTTP API, `openai` the model
provider (`anthropic`, `google` and `ollama` also exist), and `doc-processing`
the document converter (PDF, Office, HTML, Markdown, images). All extras:
[Installation](https://growgraph.github.io/ontocast/getting_started/installation/).

## Quick start

OntoCast reads its settings from environment variables:

```bash
export LLM_API_KEY=sk-...
ontocast serve
curl -X POST http://127.0.0.1:8999/process -F "file=@document.pdf" -o result.json
```

The response holds the facts and the ontology as Turtle. To process files
without a server:

```bash
ontocast process --input-path ./papers --output-dir ./out
```

To keep settings in a file, copy [`.env.example.minimal`](.env.example.minimal)
to `.env` and load it into your shell with `set -a; source .env; set +a`:
OntoCast does not read the file itself. Step by step:
[Quick start](https://growgraph.github.io/ontocast/getting_started/quickstart/).

## Your own ontologies

Put Turtle files in a directory and pass it at startup, or upload them to a
running server:

```bash
ontocast serve --ontology-dir ./my-ontologies
curl -X POST http://127.0.0.1:8999/ontologies -F "file=@my-ontology.ttl"
```

## Embed in your agent

```python
from langchain.agents import create_agent
from ontocast import Config, ToolBox, ontocast_tools

tools = await ToolBox.acreate(Config.in_memory())
await tools.initialize()

agent = create_agent(
    model,
    tools=[*ontocast_tools(tools)],
    prompt="Edit the ontology from the user's text.",
)
```

`run_unit_pipeline` processes a single passage, and `make_ontocast_node` adds
OntoCast to your own LangGraph. See [Embedding
OntoCast](https://growgraph.github.io/ontocast/guides/embedding/).

## How it works

![The OntoCast pipeline: convert, chunk, then the ontology stages and the facts stages, then serialize](https://raw.githubusercontent.com/growgraph/ontocast/refs/heads/main/docs/assets/graph.lr.png)

A document is converted to text and cut into parts. Each part updates the
ontology; the updates are normalized, consolidated and checked. Each part then
yields facts in the ontology's terms; the facts are merged, disambiguated and
validated, and the result is written to the triple store. [How OntoCast
works](https://growgraph.github.io/ontocast/concepts/).

## Documentation

| | |
|---|---|
| Getting started | [Installation](https://growgraph.github.io/ontocast/getting_started/installation/) · [Quick start](https://growgraph.github.io/ontocast/getting_started/quickstart/) |
| Concepts | [How OntoCast works](https://growgraph.github.io/ontocast/concepts/) · [Ontologies and facts](https://growgraph.github.io/ontocast/concepts/ontologies_and_facts/) |
| Guides | [Configuring OntoCast](https://growgraph.github.io/ontocast/guides/configuration/) · [Recipes](https://growgraph.github.io/ontocast/guides/recipes/) · [Validation and SHACL](https://growgraph.github.io/ontocast/guides/validation/) · [Triple stores](https://growgraph.github.io/ontocast/guides/triple_stores/) |
| Reference | [Configuration](https://growgraph.github.io/ontocast/reference/configuration/) · [HTTP API](https://growgraph.github.io/ontocast/reference/http_api/) · [Python API](https://growgraph.github.io/ontocast/reference/python/) |

Release notes: [CHANGELOG.md](CHANGELOG.md)

## Contributing

See [Contributing](https://growgraph.github.io/ontocast/contributing/). Issues
and discussion: [GitHub](https://github.com/growgraph/ontocast). Contributors
accept the [Contributor License Agreement](CLA.md) once, by commenting on
their first pull request.

## License

Apache License 2.0; see [LICENSE](LICENSE).
