Metadata-Version: 2.4
Name: graphrag-document-graph
Version: 3.1.1
Summary: Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune
Project-URL: Homepage, https://github.com/kanjani-ai-research/document-graph
Project-URL: Repository, https://github.com/kanjani-ai-research/document-graph
Author-email: Evan Erwee <evan@erwee.com>
License-Expression: MIT
License-File: LICENSE
License-File: NOTICE
Keywords: documents,etl,graph,graphrag,knowledge-graph,neptune
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Topic :: Database
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: graphrag-lexical-graph>=3.18.0
Provides-Extra: dev
Requires-Dist: black; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest-mock>=3.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff; extra == 'dev'
Description-Content-Type: text/markdown

# Document Graph

[![PyPI version](https://img.shields.io/pypi/v/graphrag-document-graph.svg)](https://pypi.org/project/graphrag-document-graph/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

**Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune.**

> **This package depends on AWS GraphRAG Toolkit (graphrag-toolkit-lexical-graph) for graph storage, vector indexing, and retrieval.**

## Installation

```bash
pip install graphrag-document-graph
```

## Dependencies

- `graphrag-toolkit-lexical-graph>=3.18.0` — AWS GraphRAG Toolkit foundation (graph storage, vector indexing, Neptune/AOSS writers)

This package does **not** depend on `bona` or `graphrag-codeproperty-graph`.

## Dependency Chain

```
graphrag-document-graph
└── graphrag-toolkit-lexical-graph>=3.18.0  (AWS foundation)
    ├── Neptune graph storage
    ├── OpenSearch Serverless vector indexing
    └── Entity resolution & retrieval
```

## Quick Start

### Python API Example

```python
from graphrag_toolkit.document_graph import PipelineExecutor, Node, Edge
from graphrag_toolkit.document_graph.graph_build.cypher_builder import CypherBuilder

# Create typed nodes from structured data
node = Node(
    id="doc-001",
    labels=["Document"],
    properties={"title": "Q4 Report", "source": "confluence", "tenant_id": "acme"}
)

# Build Cypher for Neptune ingestion
cypher, params = CypherBuilder.node_to_cypher(node, tenant_id="acme")

# Execute against Neptune via graphrag-toolkit GraphStore
graph_store.execute_query(cypher, params)
```

### Schema-Driven Pipeline

```python
from graphrag_toolkit.document_graph.schema.providers.csv_schema_provider import CSVSchemaProvider
from graphrag_toolkit.document_graph.schema.providers.schema_provider_config import SchemaProviderConfig

# Auto-discover schema from CSV
config = SchemaProviderConfig(type="csv", connection_config={"path": "data/employees.csv"})
provider = CSVSchemaProvider(config)
schema = provider.load_schema()

# Transform and load
from graphrag_toolkit.document_graph.transform.graph_transformers.row_to_node import RowToNodeTransformer
from graphrag_toolkit.document_graph.transform.transformer_provider_config import TransformerProviderConfig

transformer = RowToNodeTransformer(TransformerProviderConfig(name="r2n", args={"type": "Employee"}))
nodes = transformer.transform(records)
```

### Hybrid Search (Document Graph + Lexical Graph)

```python
from graphrag_toolkit.lexical_graph import LexicalGraphIndex, LexicalGraphQueryEngine
from graphrag_toolkit.lexical_graph.storage import GraphStoreFactory, VectorStoreFactory

# Write document-graph nodes, then index into lexical-graph for semantic search
graph_store = GraphStoreFactory.for_graph_store("neptune-db://endpoint:8182").__enter__()
vector_store = VectorStoreFactory.for_vector_store("aoss://endpoint")

graph_index = LexicalGraphIndex(graph_store, vector_store)
graph_index.extract_and_build(docs, show_progress=True)

# Semantic query across both structured and unstructured data
query_engine = LexicalGraphQueryEngine.for_traversal_based_search(graph_store, vector_store)
results = query_engine.retrieve("Who are the senior engineers?")
```

## Package Structure

```
src/graphrag_toolkit/document_graph/
├── __init__.py             # Public API: PipelineExecutor, Node, Edge, models
├── config.py               # Configuration
├── errors.py               # Custom exceptions
├── model.py                # NodeModel, EdgeModel
├── model_elements.py       # Node, Edge primitives
├── pipeline_executor.py    # Orchestrates ingest → transform → build
├── schema/                 # ETL schema model, providers, discovery
│   ├── providers/          # CSV, JSON, S3, Static, File, Glue
│   └── discovery/          # Auto-infer schema from data files
├── ingest/                 # Data ingestion (column, field, row processors)
├── transform/              # 20+ transformers
│   ├── normalizers/        # Whitespace, nulls, case, enum, timestamp
│   ├── field_transformers/ # JSON flattener, UUID gen, regex clean
│   ├── document_transformers/  # JSON to rows, text chunker, PII redactor
│   ├── filter_transformers/    # Row filter, column pruner
│   ├── graph_transformers/     # Row to node, infer edges
│   └── truncators/         # Length, field count, token limits
├── graph_build/            # Cypher generation with tenant-scoped labels
│   └── constructors/       # Node/edge Cypher builders
├── query/                  # DocumentGraphQueryEngine
├── pipeline/               # Extract (CSV, Excel, JSON, Parquet) and load
└── plugins/                # Plugin system for extensibility
```

## Integration with AWS GraphRAG Toolkit

Document Graph extends the [AWS GraphRAG Toolkit](https://github.com/awslabs/graphrag-toolkit) architecture:

```
┌─────────────────────────────────────────────────────┐
│          document-graph (this package)               │
│  Structured ETL: CSV, Excel, JSON, PDF → typed nodes│
├─────────────────────────────────────────────────────┤
│     graphrag-toolkit-lexical-graph (foundation)     │
│  GraphStore, VectorStore, Neptune/AOSS writers      │
│  LexicalGraphIndex, entity resolution, retrieval    │
└─────────────────────────────────────────────────────┘
```

## Multi-Tenancy

All operations use tenant-scoped labels for complete data isolation:

```python
node_to_cypher(node, tenant_id="acme_corp")   # → MERGE (n:`__User__acme_corp__` ...)
node_to_cypher(node, tenant_id="beta_inc")    # → MERGE (n:`__User__beta_inc__` ...)
```

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md) for development setup, testing, and PR guidelines.

## License

MIT — see [LICENSE](LICENSE) for details.

See [NOTICE](NOTICE) for third-party acknowledgments.
