Metadata-Version: 2.4
Name: de-agentic
Version: 2.1.1
Summary: AI-Powered Agentic System for Data Engineers
Author: DE Agentic Team
License-Expression: MIT
Project-URL: Homepage, https://github.com/yourusername/de-agentic
Project-URL: Documentation, https://github.com/yourusername/de-agentic/docs
Project-URL: Repository, https://github.com/yourusername/de-agentic
Keywords: data-engineering,ai-agent,llm,data-quality,etl
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pyyaml>=6.0
Requires-Dist: pydantic>=2.0
Provides-Extra: pretty
Requires-Dist: rich>=13.0; extra == "pretty"
Provides-Extra: data
Requires-Dist: duckdb>=0.9; extra == "data"
Requires-Dist: pandas>=2.0; extra == "data"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Provides-Extra: test
Requires-Dist: sqlparse>=0.4; extra == "test"
Requires-Dist: loguru>=0.7; extra == "test"
Requires-Dist: python-dotenv>=1.0; extra == "test"
Provides-Extra: legacy
Requires-Dist: typer>=0.9; extra == "legacy"
Requires-Dist: rich>=13.0; extra == "legacy"
Requires-Dist: duckdb>=0.9; extra == "legacy"
Requires-Dist: pandas>=2.0; extra == "legacy"
Requires-Dist: langchain>=0.2; extra == "legacy"
Requires-Dist: langchain-anthropic>=0.1; extra == "legacy"
Requires-Dist: langchain-openai>=0.1; extra == "legacy"
Requires-Dist: langchain-community>=0.0; extra == "legacy"
Requires-Dist: loguru>=0.7; extra == "legacy"
Requires-Dist: sqlalchemy>=2.0; extra == "legacy"
Requires-Dist: sqlparse>=0.4; extra == "legacy"
Requires-Dist: python-dotenv>=1.0; extra == "legacy"
Dynamic: license-file

# 🤖 DE Agentic - AI-Powered Data Engineering Assistant

## ▶ Start here

**Thirty seconds, no install:**

```bash
python3 de.py demo
```

Creates a sample database, profiles a file, runs a query, reads the schema and
answers a question — standard library only. No API key, no `pip install`.

Prefer a `de` command on your PATH?

```bash
python3 -m pip install -e .   # then: de demo
```

The four commands:

```bash
python3 de.py profile customers.csv                      # profile a file
python3 de.py query  "SELECT * FROM customers LIMIT 5"   # read-only SQL
python3 de.py schema                                     # tables + relationships
python3 de.py ask   "how do I find null values?"         # model, or offline help
```

Plug in a real model (llama.cpp, Ollama, OpenAI or Anthropic) whenever you want:

```bash
python3 de.py doctor      # what is reachable?
python3 de.py setup       # detect one and write .env
python3 de.py demo --ai   # full demo, step 5 answered by your model
python3 de.py tc          # end-to-end test: the model writes SQL, we run it
```

llama.cpp, Ollama and OpenAI all speak the same OpenAI-compatible
`/v1/chat/completions`, so one code path covers them. Precedence is
**environment > `.env` > `config.yaml` > defaults**.


## Cloud warehouses (ClickHouse Cloud and Snowflake)

The default remains a local database. To use a configured warehouse, put non-secret
connection settings in `config.yaml` and keep the password in an environment variable:

```yaml
database:
  type: clickhouse_cloud       # or: snowflake
  host: your-service.clickhouse.cloud
  port: 8443                   # Snowflake commonly uses 443
  database: analytics
  username: default
  password_env: CLICKHOUSE_PASSWORD
  secure: true
  # Snowflake also accepts warehouse, role and schema.
```

Then export the password and use `--db configured`:

```bash
export CLICKHOUSE_PASSWORD='...'
python3 -m pip install clickhouse-connect
python3 de.py profile --db configured
python3 de.py schema --db configured
python3 de.py query "SELECT count() FROM events" --db configured
```

For Snowflake, install `snowflake-connector-python` and set
`SNOWFLAKE_PASSWORD` (or your chosen environment-variable name). Passwords are never
stored in YAML; the YAML stores only the variable name.

📖 **Full getting-started guide: [QUICKSTART.md](QUICKSTART.md)**

> **Which docs are current.**
>
> Start with **Quick Start** and **Usage** above — they are verified against a
> real install.
>
> `docs/COMMAND_REFERENCE.md`, `docs/TESTING.md`, `docs/QUICK_TEST_REFERENCE.md`,
> `docs/DEPLOYMENT.md`, `docs/MODEL_SELECTION.md`, `docs/AGENT_MODES_OPTIONS.md`
> and `docs/FIX_CLI_OPTIONS.md` still document the legacy
> `python -m src.cli run ...` interface, which needs the optional `legacy`
> extra. They are kept for reference while that tree is migrated.
>
> `docs/MIGRATION.md` is the guide for moving off them, and
> `docs/QUALITY_REPORT.md` / `docs/USER_TESTING.md` are dated records of the
> simplification work, not living instructions.

---

A comprehensive agentic system designed to assist data engineers with daily tasks including data ingestion, modeling, quality checks, warehousing, debugging, reverse engineering, and architecture simplification.

**🆕 Now with Local LLM Support!** Run completely offline with Ollama, or use OpenAI/Anthropic for cloud-based models. Switch between models per command for cost/quality optimization.

*!Note: This is the part of README.md of Data-Agent Project, the project is covering for all topics related to data engineering. The project is still in development, but I would like to publish the documentation for reference and idea brainstorming into the to accelerating the working.*

## System Overview

![System Overview](./docs/DE_Agentic_Architecture.png)

To understand how the project is going to resolve, run the demo below:

> python ./examples/demo_local.py

> python ./examples/demo_simple.py

## 🌟 Features

### Task Categories
- **Data Ingestion**: Automated data loading from various sources (APIs, databases, files)
- **Data Modeling**: Schema design, ERD generation, normalization checks
- **Data Quality**: Profiling, validation, anomaly detection
- **Warehousing**: Pipeline generation, optimization, dbt assistance
- **Debugging**: Error analysis, log parsing, performance profiling
- **Reverse Engineering**: Schema extraction, lineage tracking, documentation generation
- **Architecture**: Pattern detection, optimization recommendations, diagram generation

### Execution Modes
- **🚀 Core Modes** (Free, Instant, No LLM Required):
  - `query`: Execute SQL queries on local database
  - `reverse`: Analyze database schema and generate documentation
  - `profile`: Profile data files with comprehensive statistics
  - `interactive`: Command-based system for database operations

- **🤖 Agent Modes** (LLM-Powered, Flexible Model Selection):
  - `debug`: Analyze errors and provide debugging guidance
  - `quality`: Data quality assessment with recommendations
  - `model`: Data modeling assistance
  - `ingest`: Data ingestion strategy and code generation
  - `warehouse`: Data warehouse architecture guidance
  - `architect`: System architecture analysis and optimization

### Mental Model Principles
- **ReAct (Reasoning + Acting)**: Combines reasoning and action for better decision-making
- **Chain of Thought**: Step-by-step reasoning for complex problems
- **Planning**: Task decomposition and strategic planning
- **Reflection**: Self-evaluation and continuous improvement
- **Memory**: Context retention across interactions

### LLM Flexibility
- **Multiple Providers**: OpenAI (GPT-4, GPT-4o-mini), Ollama (llama2, mistral, codellama), Anthropic (Claude)
- **Per-Command Model Selection**: Choose optimal model for each task using `--model` parameter
- **Cost Optimization**: Use cheap models for daily work, premium for critical tasks
- **Local-First**: Default to Ollama for free, offline operation

## 🚀 Quick Start

### Prerequisites
- Python 3.9 or newer
- Optional: [Ollama](https://ollama.ai) or another local model server
- Optional: an API key for OpenAI or Anthropic
- Optional: Docker, for the container image

### Install

```bash
pip install de-agentic
```

That is the whole thing. The base install depends only on `pyyaml` and
`pydantic`; `profile`, `query`, `schema` and `init-db` then run on the
standard library alone. Two extras are available:

```bash
pip install "de-agentic[pretty]"   # rich tables (falls back to plain text)
pip install "de-agentic[data]"     # DuckDB files, Parquet/Excel profiling
```

`de` is now on your PATH. Check it:

```bash
de --version
```

Prefer not to install anything? `python3 de.py demo` runs the same
walkthrough from a clone using only the standard library.

### Create the sample database

The commands below read `sample.sqlite`, which is generated on demand:

```bash
de init-db      # customers, products, orders, order_items
```

### Everyday commands

```bash
de profile customers.csv                       # column types, nulls, ranges
de query  "SELECT * FROM customers LIMIT 5"    # read-only SQL
de schema                                      # tables, columns, relationships
de ask    "how do I find null values?"         # answered by your model
```

`ask` works offline too: without a model it falls back to local guidance and
tells you how to configure one.

### Connect a model

```bash
de doctor      # which endpoints are reachable?
de setup       # detect one and write .env
```

Then re-run `de ask`, or try the full walkthrough with a real model:

```bash
de demo        # end to end
de demo --ai   # same, with step 5 answered by your model
```

llama.cpp, Ollama and OpenAI all speak the same OpenAI-compatible
`/v1/chat/completions`, so one code path covers them. Precedence is
**environment > `.env` > `config.yaml` > defaults**.

Per-command overrides are available when you want to switch providers
without editing configuration:

```bash
de ask "..." --provider openai --model gpt-4o-mini
```

### Using Docker

#### Quick Start with Docker Compose

```bash
# Build and start all services (app + postgres + ollama)
docker compose up -d

# Check status
docker compose ps

# View logs
docker compose logs -f de-agentic
```

#### Test the Deployment

This stack exists for the legacy agent interface, which needs the optional
`legacy` extra that the base image does not install. If you only want the
`de` commands, use the single-container form below instead.

```bash
# Interactive shell with the CLI (requires -it)
docker compose exec -it de-agentic de --help

# Create the sample database and inspect it
docker compose exec de-agentic de init-db
docker compose exec de-agentic de schema
docker compose exec de-agentic de query "SELECT * FROM customers LIMIT 5"
```

#### Single Container (No Dependencies)

The image installs this package from source, so the commands inside the
container are the same `de` commands:

```bash
# Build the image
docker build -t de-agentic .

# Standard-library walkthrough, no model required
docker run --rm de-agentic de demo

# Or use the CLI directly
docker run --rm de-agentic de schema
docker run --rm de-agentic de query "SELECT 1 AS x"

# Profile a local file (mount volume)
docker run --rm -v "$PWD/customers.csv:/data/customers.csv" \
  de-agentic de profile /data/customers.csv
```

Note: the container's default command is `de demo`, not the legacy
`python -m src.cli`, which needs optional dependencies the base image
does not install.

**📚 For complete deployment guide, see [docs/DEPLOYMENT.md](docs/DEPLOYMENT.md)**

## 📖 Usage

### Command Line Interface

#### Core modes (free, instant, no model required)

These run on the standard library and need no model server:

```bash
de init-db                                       # create sample.sqlite
de query  "SELECT COUNT(*) FROM customers"       # read-only SQL
de schema                                        # tables, columns, relationships
de profile customers.csv                         # column types, nulls, ranges
```

#### Model-backed commands

`ask` and `demo --ai` route through whichever provider is configured. Use
`de doctor` to see what is reachable and `de setup` to configure one.

```bash
de ask "which tables have no primary key?"       # answered by your model
de demo --ai                                     # full walkthrough, model answers
de tc                                            # end-to-end test of the SQL path
```

Select a provider per command without editing any config file:

```bash
de ask "..." --provider openai --model gpt-4o-mini
de ask "..." --provider ollama  --model llama3.2:1b
```

#### Legacy CLI

The earlier agent-mode interface (`run query`, `run debug`, `run reverse`,
...) is still in the tree as `src.cli`, but it is mid-migration and needs the
optional `legacy` extra:

```bash
pip install "de-agentic[legacy]"
de-agentic --help
```

It is not required for anything documented above. New code should target the
`de` commands.

#### Legacy agent modes

Available only with the `legacy` extra, via `de-agentic`. Listed for
reference; these are mid-migration and not covered by CI.

```bash
de-agentic run debug   --error="Connection timeout"
de-agentic run quality --file=data.csv --model=gpt-4o-mini
de-agentic run model   --description="E-commerce system"
de-agentic run ingest  --source="API" --target="postgres"
de-agentic run warehouse --requirements="Real-time analytics"
de-agentic run architect --description="Current pipeline"
```

Per-command `--model` works the same way here as with `de`.

### Python API

```python
from src.agents.de_agent import DEAgent
from src.tasks import DataQualityTask

# Initialize agent
agent = DEAgent()

# Run data quality check
task = DataQualityTask(
    file_path="customers.csv",
    profile=True,
    validate=True
)

result = agent.execute(task)
print(result)
```

### Available Models

| Provider | Model | Cost | Best For |
|----------|-------|------|----------|
| **OpenAI** | gpt-4o-mini | $0.15/1M tokens | Daily work, fast responses |
| | gpt-4 | $30/1M tokens | Critical issues, complex problems |
| | gpt-4-turbo | $10/1M tokens | Balance of cost/quality |
| **Ollama** | llama2 | Free | General purpose, offline |
| | mistral | Free | Better code understanding |
| | codellama | Free | Code generation, debugging |
| **Anthropic** | claude-3.5-sonnet | $3/1M tokens | Long context, analysis |
| | claude-3-opus | $15/1M tokens | Premium quality |

## 🏗️ Architecture

```
de-agentic/
├── src/
│   ├── agents/          # Agent implementations
│   ├── tasks/           # Task definitions
│   ├── skills/          # Reusable skills
│   ├── tools/           # Integration tools
│   ├── workflows/       # Workflow orchestration
│   └── utils/           # Utilities
├── config/              # Configuration files
├── examples/            # Usage examples
├── tests/               # Test suite
└── docs/                # Documentation
```

## 🔧 Configuration

### LLM Provider Configuration (.env)

```bash
# Default: Local Ollama (free, offline)
LLM_PROVIDER=ollama
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=llama2

# OpenAI (requires API key)
LLM_PROVIDER=openai
OPENAI_API_KEY=sk-proj-...
OPENAI_MODEL=gpt-4o-mini

# Anthropic (requires API key)
LLM_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...
ANTHROPIC_MODEL=claude-3.5-sonnet
```

### Agent Configuration (config/agent_config.yaml)

Edit `config/agent_config.yaml` to customize:
- Enable/disable specific tasks and skills
- Adjust mental model parameters
- Configure database connections
- Set execution limits
- Fine-tune agent behavior

### Database Setup

The `demo.db` DuckDB database includes:
- **customers** table: 10 sample customers
- **orders** table: 30 sample orders
- **products** table: 10 sample products
- Total revenue: $6,813.97

Recreate anytime with: `de init-db --force`

## 🤝 Contributing

Contributions are welcome! Please read our contributing guidelines first.

## 📚 Documentation

Comprehensive guides available in the `docs/` directory:

- **[QUICKSTART.md](QUICKSTART.md)**: 5-minute getting started guide
- **[DEPLOYMENT.md](docs/DEPLOYMENT.md)**: Complete deployment and testing guide (Docker, remote, production)
- **[MODEL_SELECTION.md](docs/MODEL_SELECTION.md)**: Complete guide to choosing and using different LLM models
- **[COMMAND_REFERENCE.md](docs/COMMAND_REFERENCE.md)**: All CLI commands with examples
- **[AGENT_MODES_OPTIONS.md](docs/AGENT_MODES_OPTIONS.md)**: Detailed comparison of all 10 execution modes
- **[DESIGN.md](DESIGN.md)**: Architecture and design principles

## 🎯 Use Cases

### Daily Data Engineering Tasks
- Query databases without remembering SQL syntax
- Profile new data files instantly
- Debug pipeline errors with AI assistance
- Generate data models from requirements

### Cost Optimization Strategy
```bash
# Development/Testing (free, local)
--model=llama2

# Daily work (cheap, cloud)
--model=gpt-4o-mini

# Critical issues (premium, best quality)
--model=gpt-4
```

### Example workflow

```bash
# 1. Profile a new file (free, instant, no model)
de profile new_data.csv

# 2. Read the schema
de schema

# 3. Run a check with a cheap model
de ask "does new_data.csv have quality issues?" --model gpt-4o-mini

# 4. Escalate to a stronger model for a hard question
de ask "how should I fix the encoding problem?" --model gpt-4

# 5. Verify with SQL (free, instant)
de query "SELECT count(*) FROM new_data"
```

## 🚀 Recent Updates

### v2.1.1 - Documentation accuracy
- ✅ Recent Updates now covers the 2.1.x line instead of stopping at v1.2.0
- ✅ Added the MIT `LICENSE` file that `pyproject.toml` and the README referenced
- ✅ Legacy reference docs are banner-marked and point at `MIGRATION.md`
- ✅ Release workflow publishes only for strict `vX.Y.Z` tags

### v2.1.0 - Packaging, CI and install docs
- ✅ GitHub Actions: lint, a 3.9-3.12 test matrix, and a build that verifies the wheel
- ✅ Tag-driven release: publish to PyPI via trusted publishing, plus a GitHub release
- ✅ The wheel now declares its dependencies; `import core.config` works after install
- ✅ `harness` data files and the `src` package ship correctly
- ✅ `de doctor` no longer crashes when no model is configured
- ✅ README and Dockerfile now match what actually installs and runs

### v1.2.0 - Model Selection Feature
- ✅ Added `--model` parameter for per-command model selection
- ✅ Auto-provider detection (gpt* → openai, claude* → anthropic)
- ✅ Cost optimization through flexible model selection
- ✅ Comprehensive documentation (docs/MODEL_SELECTION.md)

### v1.1.0 - Local Setup & CLI Enhancement
- ✅ DuckDB local database with sample data
- ✅ Ollama integration for local LLM support
- ✅ Legacy execution modes (4 core + 6 agent, behind the `legacy` extra)
- ✅ Enhanced CLI with rich terminal output
- ✅ Upgraded to Typer 0.21.0 for better compatibility

### v1.0.0 - Initial Release
- ✅ 7 task categories with modular architecture
- ✅ Mental model principles (ReAct, CoT, Planning, Reflection, Memory)
- ✅ Multiple LLM provider support
- ✅ Docker deployment support

## 📄 License

MIT License - see LICENSE file for details

## 🙏 Acknowledgments

Built with:
- **LangChain** for LLM orchestration and agent framework
- **Ollama** for local LLM deployment
- **OpenAI** & **Anthropic** for cloud LLM options
- **DuckDB** for in-memory analytics and local database
- **Typer** for CLI framework
- **Rich** for beautiful terminal output
- **Pandas** & **NumPy** for data manipulation
- **SQLAlchemy** for database connectivity
- **Great Expectations** for data quality (optional)
- **SQLGlot** for SQL parsing (optional)

## 💡 Pro Tips

1. **Start with Core Modes**: Use `query`, `reverse`, `profile`, `interactive` for free, instant results
2. **Cost Control**: Use `--model=gpt-4o-mini` for daily work, save `--model=gpt-4` for critical issues
3. **Offline Mode**: Configure `LLM_PROVIDER=ollama` in .env for completely offline operation
4. **Model Selection**: See `docs/MODEL_SELECTION.md` for detailed guidance on choosing models
5. **Agent Modes**: Read the first response from agent modes (it's usually complete) before any looping occurs

## 🐛 Troubleshooting

### Ollama Issues
```bash
# Check if Ollama is running
ollama list

# Restart Ollama service
# Windows: Restart from system tray
# Linux: systemctl restart ollama
```

### Model Not Found
```bash
# Pull the model first
ollama pull llama2
ollama pull mistral
```

### Database Issues
```bash
# Recreate the sample database
de init-db --force
```

### Query Syntax
```bash
# The SQL goes in quotes, after the subcommand
de query "SELECT * FROM customers"
```

For more troubleshooting help, see `docs/MODEL_SELECTION.md` and `docs/COMMAND_REFERENCE.md`.

## 📞 Support

- 📖 Documentation: See `docs/` directory
- 🐛 Issues: GitHub Issues
- 💬 Discussions: GitHub Discussions

---

**Ready to get started?**

```bash
pip install de-agentic
de init-db
de demo
```
# llm-based-data-engineering-agents
