Metadata-Version: 2.4
Name: tei-loop
Version: 0.1.2
Summary: Target, Evaluate, Improve: A self-improving loop for agentic systems
Author-email: Orkhan Javadli <ojavadli@gmail.com>
License: MIT
Project-URL: Homepage, https://github.com/ojavadli/tei-loop
Project-URL: Repository, https://github.com/ojavadli/tei-loop
Keywords: agents,evaluation,improvement,llm,agentic-systems,self-improving
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.0
Provides-Extra: openai
Requires-Dist: openai>=1.40; extra == "openai"
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.34; extra == "anthropic"
Provides-Extra: google
Requires-Dist: google-generativeai>=0.8; extra == "google"
Provides-Extra: all
Requires-Dist: openai>=1.40; extra == "all"
Requires-Dist: anthropic>=0.34; extra == "all"
Requires-Dist: google-generativeai>=0.8; extra == "all"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
Requires-Dist: ruff>=0.5; extra == "dev"
Dynamic: license-file

# TEI Loop

**Target, Evaluate, Improve** — a self-improving loop for agentic systems.

TEI wraps any Python agent as a black box, evaluates its output across 4 dimensions using assertion-based LLM judges, and automatically applies targeted fixes when failures are detected.

## Install

```bash
pip install tei-loop

# With your preferred LLM provider:
pip install 'tei-loop[openai]'    # OpenAI
pip install 'tei-loop[anthropic]' # Anthropic
pip install 'tei-loop[google]'    # Google Gemini
pip install 'tei-loop[all]'       # All providers
```

## Quick Start

```python
import asyncio
from tei_loop import TEILoop

# Your existing agent — any function that takes input and returns output
def my_agent(query: str) -> str:
    # ... your agent logic ...
    return result

async def main():
    loop = TEILoop(agent=my_agent)

    # Evaluate only (baseline measurement)
    result = await loop.evaluate_only("your test query")
    print(result.summary())

    # Full TEI loop (evaluate + improve + retry)
    result = await loop.run("your test query")
    print(result.summary())

    # Before/after comparison
    comparison = await loop.compare("your test query")
    print(f"Baseline: {comparison['baseline'].baseline_score:.2f}")
    print(f"With TEI: {comparison['improved'].final_score:.2f}")

asyncio.run(main())
```

## CLI

```bash
# Evaluate your agent (baseline)
tei evaluate my_agent.py --query "test input" --verbose

# Run full improvement loop
tei improve my_agent.py --query "test input" --max-retries 3

# Before/after comparison
tei compare my_agent.py --query "test input"

# Generate config file
tei init
```

TEI CLI looks for a function named `agent`, `run`, or `main` in your Python file.

## How It Works

### 1. Target
Define what success looks like. TEI evaluates across 4 dimensions:

| Dimension | What it checks |
|---|---|
| **Target Alignment** | Did the agent pursue the correct objective? |
| **Reasoning Soundness** | Was the reasoning logical and non-contradictory? |
| **Execution Accuracy** | Were the right tools called with correct parameters? |
| **Output Integrity** | Is the output complete, accurate, and consistent? |

### 2. Evaluate
TEI runs 4 LLM judges in parallel (~3-5 seconds). Each judge produces verifiable assertions, not subjective scores. Every claim is backed by evidence from the agent's output.

### 3. Improve
When a dimension fails, TEI applies the targeted fix strategy:

| Failure | Fix Strategy |
|---|---|
| Target drift | Re-anchor to original objective |
| Flawed reasoning | Regenerate plan with failure context |
| Execution errors | Correct tool calls and parameters |
| Output issues | Repair factual errors and fill gaps |

The loop retries automatically. Each cycle is sharper because the last was diagnosed.

## Two Modes

**Runtime mode** (default): Per-query, 1-3 retries, fixes individual failures in seconds.

**Development mode**: Across many queries, proposes permanent prompt improvements.

```python
# Development mode
dev_results = await loop.develop(
    queries=["query1", "query2", "query3", ...],
    max_iterations=50,
)
print(f"Avg improvement: {dev_results['avg_improvement']:+.2f}")
```

## Configuration

TEI auto-detects your LLM provider from environment variables (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GOOGLE_API_KEY`). No new accounts or API keys needed.

```python
loop = TEILoop(
    agent=my_agent,
    eval_llm="gpt-5.2",          # Smartest for evaluation
    improve_llm="gpt-5-mini",  # Cost-effective for fixes
    max_retries=3,
    verbose=True,
)
```

Or via `.tei.yaml`:

```bash
tei init  # generates config file
```

## Cost

| Scenario | Cost |
|---|---|
| Agent passes all dimensions | ~$0.005 |
| One improvement cycle | ~$0.025 |
| Full 3-retry loop | ~$0.07 |

TEI shows the cost estimate before running.

## Works With

TEI wraps any Python callable. No framework lock-in:

- **LangGraph** agents
- **CrewAI** crews
- **Custom Python** functions
- **FastAPI** endpoints
- **Any callable** that takes input and returns output

## License

MIT
