Metadata-Version: 2.4
Name: ai-team
Version: 1.0.0rc1
Summary: Experimental, CLI-first autonomous software-engineering orchestration platform with planning, sandboxed execution, validation, Git automation, approvals, and deployment workflows.
License: MIT License
        
        Copyright (c) 2026 Jayesh Patil
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACTSpring, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/Jayesh01323/AI_Team
Project-URL: Repository, https://github.com/Jayesh01323/AI_Team
Project-URL: Documentation, https://github.com/Jayesh01323/AI_Team/blob/master/docs/README.md
Project-URL: Changelog, https://github.com/Jayesh01323/AI_Team/blob/master/CHANGELOG.md
Project-URL: Issues, https://github.com/Jayesh01323/AI_Team/issues
Keywords: ai,autonomous-agents,software-engineering,orchestration,llm,code-generation
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Topic :: Software Development :: Code Generators
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click>=8.0.0
Requires-Dist: fastapi>=0.100.0
Requires-Dist: openai>=1.0.0
Requires-Dist: pydantic>=2.0.0
Requires-Dist: python-dotenv>=1.0.0
Requires-Dist: uvicorn>=0.23.0
Requires-Dist: websockets>=11.0
Provides-Extra: dev
Requires-Dist: httpx>=0.24.0; extra == "dev"
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=5.0.0; extra == "dev"
Requires-Dist: sqlalchemy>=2.0.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Requires-Dist: ruff>=0.1.0; extra == "dev"
Requires-Dist: bandit>=1.7.0; extra == "dev"
Dynamic: license-file

<div align="center">

# AI_Team

**CLI-first autonomous software-engineering orchestration platform**

From a plain-language idea to a planned, executed, validated, and deployable project —
with sandboxed execution, Git automation, human approval, and a live developer control panel.

[![Python](https://img.shields.io/badge/python-3.11%2B-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Release](https://img.shields.io/badge/release-v1.0.0--rc1-orange.svg)](docs/releases/v1.0.0-rc1.md)
[![Status](https://img.shields.io/badge/status-experimental%20%2F%20proof--of--concept-lightgrey.svg)](#release-status)
[![Code style: Ruff](https://img.shields.io/badge/code%20style-ruff-000000.svg)](https://github.com/astral-sh/ruff)
[![CI](https://github.com/Jayesh01323/AI_Team/actions/workflows/ci.yml/badge.svg)](https://github.com/Jayesh01323/AI_Team/actions/workflows/ci.yml)

[Quick Start](#quick-start) · [Documentation](docs/README.md) · [Architecture](#architecture) · [M14](docs/m14/) · [M15](docs/m15/) · [Frontend](docs/frontend.md)

</div>

> **Experimental proof-of-concept.** This repository is a research project, not a finished
> product or enterprise platform. See [Release Status](#release-status) and
> [Known Limitations](#known-limitations).

---

## Workflow at a glance

```mermaid
flowchart LR
    Idea(["Idea"]) --> Brain["Engineering Brain<br/><i>6-stage LLM pipeline</i>"]
    Brain --> Plan["Plan<br/><i>requirements · PRD · spec<br/>architecture · tasks</i>"]
    Plan --> Prep["Prepare<br/><i>scaffold + init Git</i>"]
    Prep --> Exec["Execute<br/><i>adapter in ExecutionSandbox</i>"]
    Exec --> Val["Validate<br/><i>Ruff + Pytest</i>"]
    Val -->|fail| Rec["Recover<br/><i>self-healing retry</i>"]
    Rec --> Exec
    Val -->|pass| Git["Git<br/><i>branch + atomic commit</i>"]
    Git --> Appr["Approve<br/><i>HITL gate</i>"]
    Appr --> Deploy["Deploy<br/><i>manifest · validate · probe</i>"]
    Deploy --> Dash["Observe<br/><i>REST · WebSocket · React UI</i>"]

    classDef auto fill:#eef2ff,stroke:#6366f1,color:#1e1b4b;
    classDef gate fill:#fff7ed,stroke:#f59e0b,color:#7c2d12;
    class Brain,Plan,Prep,Exec,Val,Rec auto;
    class Appr gate;
```


## What is AI_Team?

AI_Team is a Python platform that turns a high-level idea into a machine-readable engineering
plan and then attempts to implement, validate, and package it autonomously. It is organised
around three concerns:

- **The Engineering Brain** — a six-stage LLM pipeline that produces requirements, a PRD, a
  project specification, an architecture, and a task plan as versioned Markdown/JSON artifacts.
- **The Execution Engine** — a provider-agnostic framework that dispatches tasks to coding-agent
  adapters inside a confined local sandbox, tracks workspace diffs, validates the result with
  Ruff and Pytest, and retries through a self-healing loop.
- **The Control Panel** — a FastAPI REST API with WebSocket event streaming, plus a React UI for
  projects, live runs, approvals, deployments, artifacts, and health.

## Why it exists

Producing a complete software project from an idea requires many coordinated engineering steps
that are usually done manually and are rarely auditable. AI_Team explores whether that chain can
be made **structured, inspectable, and reproducible**: every stage emits typed artifacts, every
execution is sandboxed and validated, every run is traced with correlation IDs, and automation
is gated by human approval where it matters.

## What V1 actually does

V1 is a working, tested proof-of-concept. Concretely, it can:

1. Turn an idea into requirements, a PRD, a specification, an architecture, and a task plan.
2. Generate a project repository scaffold from that plan.
3. Execute tasks through a coding-agent adapter inside a sandboxed workspace.
4. Run Ruff and Pytest validation against the result, with a bounded self-healing retry loop.
5. Record the work in Git — task branches, atomic commits with structured metadata, and a
   pull-request summary artifact.
6. Gate key stages behind a human-in-the-loop approval lifecycle.
7. Generate and validate container deployment manifests (Dockerfile / compose / deployment
   metadata) and probe service health where an environment is available.
8. Observe and drive all of the above through the dashboard API, WebSocket events, and React UI.

See [Known Limitations](#known-limitations) for what is deliberately **not** included.

## Architecture

The system is organised into strict, unidirectional layers. The dependency direction is
enforced at test time by AST-based boundary checks (`architecture_rules.py`).

```text
CLI (cli/)
  └── Application services (app/)
        ├── Pipeline engine (pipeline/, brain/)
        │     └── Providers (providers/)
        └── Execution engine (execution/)
              ├── Adapters (execution/adapters/)
              ├── Sandbox (execution/sandbox.py)
              ├── Validation (execution/validation/)
              └── Recovery (execution/recovery/)
Dashboard API (dashboard/) ── services ── reliability/state (core/reliability/)
Core infrastructure (core/): config, logging, exceptions, reliability, deployment, benchmark
```

| Layer | Directory | Responsibility |
| :--- | :--- | :--- |
| CLI | `cli/` | Click entry points (`init`, `analyze`, `pipeline`, `run`, …) |
| Application | `app/` | Service façade between CLI and engines |
| Brain | `brain/`, `pipeline/` | LLM stages, orchestration, artifact export |
| Providers | `providers/` | LLM provider implementations + factory |
| Execution | `execution/` | Adapters, sandbox, diff tracking, validation, recovery, Git |
| Deployment | `core/deployment/` | Manifest generation, container validation, health probing |
| Dashboard | `dashboard/` | FastAPI routers, services, WebSocket, React frontend |
| Core | `core/` | Config, logging, exceptions, reliability, benchmark/evaluation |

Architecture decisions are recorded as ADRs in [`docs/adr/`](docs/adr/) (layering, provider
registry, adapter pattern, validation pipeline, structured logging, boundary tests).
See the [documentation index](docs/README.md) for the full map.

## Core capabilities

Status legend: **Implemented** (works and is tested) · **Partial** (works, but with a bounded
scope) · **Scaffolded** (interface + lifecycle only, no live integration) ·
**Environment-dependent** (real behavior requires external services) · **Not in V1**.

| Area | What V1 provides | Status |
| :--- | :--- | :--- |
| Engineering Brain | 6-stage pipeline: idea → requirements → PRD → spec → architecture → tasks | Implemented |
| Artifacts | Markdown + JSON outputs (`requirements.md`, `PRD.md`, `architecture.json`, `task_plan.json`, …) | Implemented |
| LLM providers | OpenAI, Google Gemini, NVIDIA NIM, and `auto` fallback routing | Implemented |
| Execution adapters | OpenHands adapter; `live` / `live_llm` LLM-driven adapter | Implemented |
| Adapter registry | Pluggable `ExecutionAdapter` contract + capability validation | Implemented |
| Execution sandbox | Process-level confinement, credential isolation, timeouts, process-tree termination, redaction | Implemented |
| Validation | Ruff + Pytest, sequential and parallel | Implemented |
| Orchestration | End-to-end autonomous orchestrator with stage events | Implemented |
| Self-healing | Validation-log-driven retry / repair loop | Implemented |
| Git automation | Init, task branches, atomic commits with `.ai/commits/` metadata, merge, PR-summary artifact | Implemented |
| HITL approvals | PENDING gates, approve/reject by ID, project-scoped listing | Implemented |
| Dashboard API | FastAPI REST endpoints for projects, runs, approvals, deployments, health | Implemented |
| Event streaming | WebSocket stage/task/approval/deployment events | Implemented |
| React UI | Projects, project detail, active execution, approvals, health | Implemented |
| Reliability | Checkpointing, state store, schema/planning-artifact locks | Implemented |
| Telemetry | Structured JSONL logging with correlation IDs | Implemented |
| Security | Sandbox isolation, path-traversal protections, CORS pinning, regression tests | Implemented |
| Quality tooling | M9–M11 benchmark, evaluation, and release-engineering modules | Implemented |
| Deployment (M15) | Manifest + Dockerfile/compose generation, validation, health probing, records/logs | Partial — container build/probe are environment-dependent |
| Live LLM execution | Real provider-backed generation and task implementation | Environment-dependent — requires provider credentials |
| Other adapters | Claude, Cursor, Devin, VS Code, Codex, Antigravity | Scaffolded |
| Anthropic provider | `AnthropicProvider` | Scaffolded (declared stub) |
| Per-task container isolation | Ephemeral container sandbox per task | Not in V1 |
| GitHub PR API submission | Creating pull requests via the GitHub API | Not in V1 — a local PR-summary artifact is generated |
| Multi-agent peer review | Security Auditor / Code Reviewer agents | Not in V1 |


## How it works

1. **Planning** — `PipelineEngine` runs the registered brain stages and writes artifacts into
   `projects/<slug>/`. Each stage is an LLM call through the provider abstraction.
2. **Preparation** — the orchestrator scaffolds a repository from the plan and initialises Git.
3. **Execution** — each `ExecutionTask` is dispatched to an adapter inside an
   `ExecutionSandbox`; the workspace is an isolated copy, and file changes are tracked via diffs.
4. **Validation** — Ruff and Pytest run against the workspace through the sandbox.
5. **Recovery** — on failure, the self-healing engine builds a repair prompt and retries up to
   the configured budget.
6. **Git** — successful tasks are committed atomically on a task branch with structured metadata.
7. **Approval** — completed planning/architecture stages create PENDING approval gates.
8. **Deployment** — deployment runs generate manifests, validate them, probe health, and record
   lifecycle logs.
9. **Observability** — every transition is emitted as an event over WebSocket; checkpoints
   persist to `core/reliability` state.

## Quick Start

### Prerequisites

- **Python** 3.11 or newer
- **Git** (on `PATH`)
- *Optional:* Node.js/npm to run the React dashboard
- *Optional:* Docker for real container build/validation in the deployment lifecycle
- *Optional:* an LLM provider API key for live provider-backed generation

### 1. Install

**From PyPI (Recommended for users):**
*(Note: After the first PyPI release, this will be the standard installation method)*

```bash
pip install ai-team
```

**From Source (For developers):**

```bash
git clone https://github.com/Jayesh01323/AI_Team.git
cd AI_Team
python -m venv venv && source venv/bin/activate   # Windows: venv\Scripts\activate
pip install -e ".[dev]"                            # runtime + dev/test dependencies
```

### 2. Configure environment (only needed for provider-backed commands)

```bash
cp .env.example .env
```

```env
AI_PROVIDER=gemini        # or openai / nvidia / auto
GEMINI_API_KEY=your-key
```

### 3. First CLI command (no credentials required)

```bash
ai-team --help
ai-team init "A simple test project"
```

### 4. Start the backend API

```bash
python -m uvicorn dashboard.api:app --host 127.0.0.1 --port 8000
```

Health check: <http://127.0.0.1:8000/api/health>

### 5. Start the dashboard UI (optional, separate terminal)

```bash
cd dashboard/frontend
npm install
npm run dev        # Vite dev server on http://localhost:3000, proxying /api and /ws to :8000
```

## CLI

| Command | Purpose | Provider credentials? |
| :--- | :--- | :---: |
| `ai-team --help` | Show all commands and options | No |
| `ai-team init "<idea>"` | Create a project folder with placeholder templates | No |
| `ai-team test-provider` | Verify the configured LLM provider connection | **Yes** |
| `ai-team analyze "<idea>"` | Run the Idea Analysis stage | **Yes** |
| `ai-team generate "<idea>"` | Analyze and generate `requirements.md` | **Yes** |
| `ai-team pipeline "<idea>"` | Run the full 6-stage Engineering Brain pipeline | **Yes** |
| `ai-team scaffold "<idea>"` | Run the pipeline and generate a physical repo structure | **Yes** |
| `ai-team run "<idea>" --provider openhands` | Run the autonomous pipeline with execution + validation | **Yes** for live providers |

## Dashboard

The backend is a FastAPI app exported as `dashboard.api:app`.

**Backend** (port **8000**):

```bash
python -m uvicorn dashboard.api:app --host 127.0.0.1 --port 8000
```

**Frontend** (port **3000**): see [`docs/frontend.md`](docs/frontend.md).

```bash
cd dashboard/frontend
npm install
npm run dev
```

Key REST routes (all under `/api`): `GET /health`, `GET|POST /projects`,
`GET /projects/{id}`, `GET /projects/{id}/artifacts/{artifact_name}`,
`POST /projects/{id}/execute`, `POST /projects/{id}/deploy`,
`GET /runs`, `GET /runs/{id}`, `GET /approvals`, `POST /approvals/{id}/approve|reject`,
`GET /deployments`, `GET /deployments/{id}`, `GET /deployments/{id}/logs`.
Live events stream over `WS /ws/events`.

## Example workflow

A factual end-to-end run (provider credentials required for the planning step):

```bash
# 1. Plan + generate artifacts (requires a provider key)
ai-team pipeline "A SaaS platform for analyzing resumes"
#    -> projects/build-a-saas-resume-analyzer/{requirements.md,PRD.md,architecture.json,...}

# 2. Scaffold a repository from the plan
ai-team scaffold "A SaaS platform for analyzing resumes"

# 3. Run the autonomous pipeline (planning -> execution -> validation -> git)
ai-team run "A SaaS platform for analyzing resumes" --provider openhands

# 4. Or drive everything from the dashboard API
python -m uvicorn dashboard.api:app --port 8000
curl -X POST http://127.0.0.1:8000/api/projects \
  -H 'Content-Type: application/json' -d '{"name":"demo","raw_idea":"A todo app"}'
```

## Release Status

- **Version:** `v1.0.0-rc1` (release candidate)
- **Release commit:** `bd26c8d`
- **Branch:** `master`
- **Nature:** experimental research / proof-of-concept

Release notes: [`docs/releases/v1.0.0-rc1.md`](docs/releases/v1.0.0-rc1.md).

## Verification

The suite is credential-hermetic (live cases skip rather than fail) and organised by milestone.

```bash
python -m pytest -q                            # backend test suite
python -m ruff check core/ dashboard/ tests/   # lint
python run_m15_capstone.py                     # M15 deployment lifecycle capstone
```

- **Test suite:** ~825+ backend tests plus module-level tests under `brain/*/tests`.
- **Ruff:** clean on `core/`, `dashboard/`, `tests/`.
- **M15 capstone:** passes end-to-end (manifest generation → container validation → Git → dashboard lifecycle).
- **Security/approval regressions:** pass (artifact traversal, deployment-ID traversal, CORS pinning, approval lifecycle).
- **CI:** `.github/workflows/ci.yml` runs Ruff + Pytest on **Windows** (Python 3.11).

> **Known test-harness issue:** the M14 capstone asserts `pull_request.md` while the
> implementation writes `PULL_REQUEST.md`. This passes on case-insensitive filesystems
> (matching the current Windows CI) and fails on case-sensitive ones. It is a
> portability/test-harness mismatch, not a product defect.

## Known Limitations

- **Process-level sandboxing**, not per-task container isolation.
- **One implemented live execution backend** (OpenHands) plus the `live` LLM adapter; the
  remaining six adapters are scaffolds.
- **Anthropic provider** is a declared stub.
- **Live LLM execution requires provider credentials**; without them the pipeline falls back.
- **Deployment container build and health probing are environment-dependent** (Docker/HTTP) and
  otherwise simulated.
- **GitHub API PR submission is not implemented** — V1 generates a local pull-request summary artifact.
- **Multi-agent role peer review is not implemented.**
- **CI is Windows-only**; frontend Vitest tests are not run by CI.
- **No Python dependency lockfile** is committed.
- Several `projects/*/context.json` runtime-state files are tracked and churn when the suite runs.

## Documentation

Start at the [documentation index](docs/README.md).

| Topic | Location |
| :--- | :--- |
| Architecture decisions | [`docs/adr/`](docs/adr/) |
| Core architecture | [`docs/core/ARCHITECTURE.md`](docs/core/ARCHITECTURE.md) |
| Dashboard / M12 | [`docs/m12/`](docs/m12/) |
| Sandbox, Git & autonomous execution (M14) | [`docs/m14/`](docs/m14/) |
| Deployment lifecycle (M15) | [`docs/m15/`](docs/m15/) |
| Frontend | [`docs/frontend.md`](docs/frontend.md) |
| Security policy | [`SECURITY.md`](SECURITY.md) |
| Release notes | [`docs/releases/`](docs/releases/) |
| Contributing | [`CONTRIBUTING.md`](CONTRIBUTING.md) |

## Contributing

See [`CONTRIBUTING.md`](CONTRIBUTING.md) and the [`CODE_OF_CONDUCT.md`](CODE_OF_CONDUCT.md).

## Security

See [`SECURITY.md`](SECURITY.md) for the security policy and vulnerability reporting.

## License

MIT — see [`LICENSE`](LICENSE).

---

<div align="center">

**AI_Team** · experimental V1 release candidate `v1.0.0-rc1` (`bd26c8d`)

Built by [Jayesh Patil](https://github.com/Jayesh01323) · [LinkedIn](https://www.linkedin.com/in/jayesh-patil-5a2a082b5)

[Back to top](#ai_team) · [Documentation](docs/README.md) · [Release notes](docs/releases/v1.0.0-rc1.md)

</div>
