Metadata-Version: 2.4
Name: graphrag-document-graph
Version: 0.1.0
Summary: Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune
Project-URL: Homepage, https://github.com/RW-Lab/document-graph
Project-URL: Repository, https://github.com/RW-Lab/document-graph
Author-email: Evan Erwee <evan@erwee.com>
License-Expression: MIT
License-File: LICENSE
License-File: NOTICE
Keywords: documents,etl,graph,graphrag,knowledge-graph,neptune
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Topic :: Database
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: boto3>=1.26.0
Requires-Dist: networkx>=3.0
Requires-Dist: pandas>=2.0.0
Requires-Dist: pluggy>=1.0.0
Requires-Dist: pydantic>=2.0.0
Provides-Extra: dev
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest-mock>=3.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Provides-Extra: graphrag
Requires-Dist: graphrag-toolkit-lexical-graph>=3.18.0; extra == 'graphrag'
Provides-Extra: neo4j
Requires-Dist: neo4j>=5.0.0; extra == 'neo4j'
Provides-Extra: neptune
Requires-Dist: boto3>=1.26.0; extra == 'neptune'
Description-Content-Type: text/markdown

# Document Graph

Document knowledge graphs — extract structured data from documents (Confluence, PDFs, CSV, JSON, Excel) into Neptune.

> **This package depends on AWS GraphRAG Toolkit (graphrag-toolkit-lexical-graph) for graph storage, vector indexing, and retrieval.**

## Quick Start

### Install

```bash
pip install document-graph                    # core
pip install document-graph[graphrag]          # with lexical-graph integration
```

### Python API Example

```python
from graphrag_toolkit.document_graph import PipelineExecutor, Node, Edge
from graphrag_toolkit.document_graph.graph_build.cypher_builder import CypherBuilder

# Create typed nodes from structured data
node = Node(
    id="doc-001",
    labels=["Document"],
    properties={"title": "Q4 Report", "source": "confluence", "tenant_id": "acme"}
)

# Build Cypher for Neptune ingestion
cypher, params = CypherBuilder.node_to_cypher(node, tenant_id="acme")

# Execute against Neptune via graphrag-toolkit GraphStore
graph_store.execute_query(cypher, params)
```

### Schema-Driven Pipeline

```python
from graphrag_toolkit.document_graph.schema.providers.csv_schema_provider import CSVSchemaProvider
from graphrag_toolkit.document_graph.schema.providers.schema_provider_config import SchemaProviderConfig

# Auto-discover schema from CSV
config = SchemaProviderConfig(type="csv", connection_config={"path": "data/employees.csv"})
provider = CSVSchemaProvider(config)
schema = provider.load_schema()

# Transform and load
from graphrag_toolkit.document_graph.transform.graph_transformers.row_to_node import RowToNodeTransformer
from graphrag_toolkit.document_graph.transform.transformer_provider_config import TransformerProviderConfig

transformer = RowToNodeTransformer(TransformerProviderConfig(name="r2n", args={"type": "Employee"}))
nodes = transformer.transform(records)
```

### Hybrid Search (Document Graph + Lexical Graph)

```python
from graphrag_toolkit.lexical_graph import LexicalGraphIndex, LexicalGraphQueryEngine
from graphrag_toolkit.lexical_graph.storage import GraphStoreFactory, VectorStoreFactory
from llama_index.core.schema import Document

# Write document-graph nodes, then index into lexical-graph for semantic search
graph_store = GraphStoreFactory.for_graph_store("neptune-db://endpoint:8182").__enter__()
vector_store = VectorStoreFactory.for_vector_store("aoss://endpoint")

graph_index = LexicalGraphIndex(graph_store, vector_store)
graph_index.extract_and_build(docs, show_progress=True)

# Semantic query across both structured and unstructured data
query_engine = LexicalGraphQueryEngine.for_traversal_based_search(graph_store, vector_store)
results = query_engine.retrieve("Who are the senior engineers?")
```

## Package Structure

```
src/graphrag_toolkit/document_graph/
├── __init__.py             # Public API: PipelineExecutor, Node, Edge, models
├── config.py               # Configuration
├── errors.py               # Custom exceptions
├── model.py                # NodeModel, EdgeModel
├── model_elements.py       # Node, Edge primitives
├── pipeline_executor.py    # Orchestrates ingest → transform → build
├── schema/                 # ETL schema model, providers, discovery
│   ├── providers/          # CSV, JSON, S3, Static, File, Glue
│   └── discovery/          # Auto-infer schema from data files
├── ingest/                 # Data ingestion (column, field, row processors)
├── transform/              # 20+ transformers
│   ├── normalizers/        # Whitespace, nulls, case, enum, timestamp
│   ├── field_transformers/ # JSON flattener, UUID gen, regex clean
│   ├── document_transformers/  # JSON to rows, text chunker, PII redactor
│   ├── filter_transformers/    # Row filter, column pruner
│   ├── graph_transformers/     # Row to node, infer edges
│   └── truncators/         # Length, field count, token limits
├── graph_build/            # Cypher generation with tenant-scoped labels
│   └── constructors/       # Node/edge Cypher builders
├── query/                  # DocumentGraphQueryEngine
├── pipeline/               # Extract (CSV, Excel, JSON, Parquet) and load
└── plugins/                # Plugin system for extensibility
```

## Integration

### With AWS GraphRAG Toolkit

Document Graph extends the [AWS GraphRAG Toolkit](https://github.com/awslabs/graphrag-toolkit) architecture:

```
┌─────────────────────────────────────────────────────┐
│          document-graph (this package)               │
│  Structured ETL: CSV, Excel, JSON, PDF → typed nodes│
├─────────────────────────────────────────────────────┤
│     graphrag-toolkit-lexical-graph (foundation)     │
│  GraphStore, VectorStore, Neptune/AOSS writers      │
│  LexicalGraphIndex, entity resolution, retrieval    │
└─────────────────────────────────────────────────────┘
```

### With Neptune & OpenSearch Serverless

Deploy infrastructure via the [graphrag-toolkit CloudFormation templates](https://github.com/awslabs/graphrag-toolkit/tree/main/examples/lexical-graph/cloudformation-templates) which provisions:
- Amazon Neptune (graph database)
- Amazon OpenSearch Serverless (vector search)
- SageMaker notebook instance

### Multi-Tenancy

All operations use tenant-scoped labels for complete data isolation:

```python
# Tenant A's data is completely isolated from Tenant B
node_to_cypher(node, tenant_id="acme_corp")   # → MERGE (n:`__User__acme_corp__` ...)
node_to_cypher(node, tenant_id="beta_inc")    # → MERGE (n:`__User__beta_inc__` ...)
```

## Requirements

- Python >= 3.11
- `pydantic >= 2.0`
- `pandas >= 2.0`
- `boto3 >= 1.26`
- Optional: `graphrag-toolkit-lexical-graph >= 3.18.0` (for hybrid search)

## License

MIT — see [LICENSE](LICENSE) for details.

See [NOTICE](NOTICE) for third-party acknowledgments.
