Metadata-Version: 2.4
Name: docHandler4AISDK
Version: 1.0.2
Summary: Professional Python SDK for DocHandler4AI API supporting OCR, PDF processing, and Indexing.
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Requires-Python: >=3.8
Description-Content-Type: text/markdown
Requires-Dist: httpx>=0.24.0
Requires-Dist: pydantic>=2.0.0

# docHandler4AI Python SDK

The **docHandler4AI** Python SDK provides a high-level, asynchronous interface to the docHandler4AI server. It enables advanced PDF processing, multimodal vision analysis, and robust vector indexing for both text and images.

## Table of Contents
- [Installation](#installation)
- [Quick Start](#quick-start)
- [PDF Processing](#pdf-processing)
- [Vision Analysis](#vision-analysis)
- [Document Indexing](#document-indexing)
  - [Searching with Metadata Filters](#searching-with-metadata-filters)
  - [Self-Query Search](#self-query-search)
- [Image Indexing](#image-indexing)
  - [Multimodal Search](#multimodal-search)

---

## Model Parameters

The SDK provides typed schemas for `model_params` to help you configure specific AI models.

### Global Parameters
Applies to most models (OpenAI, Gemini, etc.):
```python
from dochandler4ai_sdk import GlobalModelParams

params = GlobalModelParams(temperature=0.7, max_tokens=1000)
result = await vision.describe(
    image_url="...",
    model_params=params
)
```

### DeepSeek Specific Parameters
DeepSeek models support additional fields like "thinking" control:

```python
from dochandler4ai_sdk import DeepSeekVisionParams

# Using the specialized schema
params = DeepSeekVisionParams(temperature=0.0, thinking={"type": "disabled"})

result = await vision.describe(
    image_url="...",
    model_name="deepseek-v4-flash-vision-exp",
    model_params=params
)
```

---

## Installation

```bash
pip install docHandler4AISDK
```

---

## Quick Start

Initialize the main client facade:

```python
import asyncio
from dochandler4ai_sdk import DocHandler4AI

async def main():
    sdk = DocHandler4AI(
        base_url="http://localhost:8000",
        api_key="your_api_key"
    )
    
    # Use the SDK...
    
asyncio.run(main())
```

---

## PDF Processing

Extract structured content from PDFs using Vision-based AI.

```python
pdf_processor = sdk.get_pdf_process()

# 1. Process PDF to Markdown (optionally with chunks)
result = await pdf_processor.to_markdown(
    pdf_url="https://example.com/document.pdf",
    return_chunks=True
)
print(result["markdown"])

# 2. Process PDF directly to Chunks
result = await pdf_processor.to_chunks(
    pdf_url="https://example.com/document.pdf"
)
for chunk in result.get("chunks", []):
    print(chunk["content"])
```

---

## Vision Analysis

Perform OCR, generate captions, or analyze specific PDF pages.

```python
vision = sdk.get_vision_process()

# 1. Analyze image (Graph, Diagram, General content)
analysis = await vision.describe(
    image_url="https://example.com/diagram.png"
)
print(analysis["content"])

# 2. Extract text (OCR)
result = await vision.ocr(
    image_url="https://example.com/text_image.png"
)
print(result["content"])

# 3. PDF Page to Markdown
# Returns a single markdown with embedded image descriptions: 
# <image alt="" description="..."/>
result = await vision.to_markdown(
    image_url="https://example.com/pdf_page_render.png"
)
print(result["content"])
```

---

## Document Indexing

Manage vector collections for text documents with metadata filtering.

### The Importance of `doc_id`
When indexing documents, the `doc_id` is a unique identifier for a **logical document** (e.g., a specific PDF file). 
*   **Logical Grouping**: A single PDF might be split into 50 chunks. All 50 chunks must share the same `doc_id`.
*   **Automatic Updates**: If you upsert chunks with a `doc_id` that already exists in the collection, the server will **automatically replace** the old chunks with the new ones. This ensures you don't have duplicate content for the same document.
*   **Deletion**: You can delete an entire document and all its associated chunks in one call using its `doc_id`.

```python
from dochandler4ai_sdk import DocumentChunk

# Get document collection
docs = sdk.get_documents_collection(
    collection_id="my_docs",
    embedding_model="openai/text-embedding-3-small/1536" # Default for this collection
)

# Upsert documents
await docs.upsert(chunks=[
    DocumentChunk(
        doc_id="doc_001",
        page_content="Artificial Intelligence is transforming industries...",
        metadata={"category": "technology", "author": "Alice"}
    )
])

# Count documents
count = await docs.count(where={"category": "technology"})
print(f"Total technology documents: {count['total']}")

# Get sample documents
samples = await docs.get(k=5, where={"author": "Alice"})
```

### Searching with Metadata Filters

The SDK supports complex metadata filtering using MongoDB-like operators ($or, $in, $gt, $lt, etc.).

```python
from dochandler4ai_sdk import DocumentSearchRequest

req = DocumentSearchRequest(
    query="How is AI changing the world?",
    k=3,
    filters={
        "$or": [
            {"category": "technology"},
            {"tags": {"$in": ["AI", "ML"]}}
        ]
    }
)

results = await docs.search(req)
```

### Self-Query Search

Let the LLM automatically translate natural language into structured filters.

```python
req = DocumentSearchRequest(
    query="Show me documents by Alice about technology written after 2023",
    k=5,
    search_type="self_query",
    metadata_field_info=[
        {"name": "author", "description": "The author of the document", "type": "string"},
        {"name": "category", "description": "The document category", "type": "string"},
        {"name": "year", "description": "Year of publication", "type": "integer"}
    ]
)
results = await docs.search(req)
```

---

## Image Indexing

Multi-modal vector storage for images.

```python
from dochandler4ai_sdk import ImageItem

images = sdk.get_images_collection(
    collection_id="product_catalog",
    embedding_model="nvidia/llama-nemotron-embed-vl-1b-v2/2048"
)

# Upsert images
await images.upsert(images=[
    ImageItem(
        image_id="img_101",
        image_url="https://example.com/headphone.jpg",
        metadata={"category": "headphone", "brand": "Sony"}
    )
])

# Count images
count = await images.count(where={"category": "headphone"})
```

### Multimodal Search

Search images using either a text query (Text-to-Image) or another image (Image-to-Image).

```python
from dochandler4ai_sdk import ImageSearchRequest

# 1. Text-to-Image Search
req = ImageSearchRequest(
    query="blue wireless headphones",
    where={"brand": "Sony"},
    score_threshold=0.7
)
results = await images.search(req)

# 2. Image-to-Image Search
req = ImageSearchRequest(
    image_url="https://example.com/reference_image.jpg",
    k=5
)
results = await images.search(req)
```
