Metadata-Version: 2.5
Name: dreamrsi
Version: 0.2.0a4
Summary: Independent, unofficial Dream-RSI-inspired Python SDK for exploration-policy replay and search (alpha).
Project-URL: Homepage, https://github.com/TheAstrayDev/dream-rsi-sdk
Project-URL: Documentation, https://github.com/TheAstrayDev/dream-rsi-sdk#readme
Author: TheAstrayDev
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: agents,ai,dream-rsi,exploration,recursive-self-improvement
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.11
Provides-Extra: dev
Requires-Dist: hypothesis>=6.0; extra == 'dev'
Requires-Dist: pyright>=1.1; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: langchain
Requires-Dist: langchain-core<2,>=1.6.3; extra == 'langchain'
Description-Content-Type: text/markdown

# Dream-RSI SDK · Alpha

**Bring your agent. Record its search. Evolve executable exploration policies.**

An independent, unofficial Python SDK inspired by Dream-RSI research. Maintained by
**TheAstrayDev**, who is not a Google or Google DeepMind employee. This is a personal
research initiative, not a commercial Google development, official SDK or endorsed product.

Python 3.11+ · zero core third-party dependencies · Apache-2.0 · version **0.2.0a4**.

## Install

```bash
python -m pip install --upgrade dreamrsi==0.2.0a4
```

The package includes the Python API and the `dreamrsi` command. Publishing policy
packages to GitHub additionally requires GitHub CLI (`gh`).

## New in 0.2.0a4

The replay optimizer can synthesize a shorter `PrefixPolicy` without model
requests: preserve the incumbent's original decisions and complete batches
until every training world's full raw quality has been reached, then stop.
Nested source policies, checkpoints and portable bundles are supported.
Independent holdout validation is still required; recorded quality does not
guarantee generalization. The original replay objective remains unchanged.

Disable proposals with `DreamRSIConfig(optimizer_prefix_search=False)`.
Developer eligibility now accounts for root-branch ordering when bounding
recorded probe cost. Prefix bundles need 0.2.0a4 or newer; existing trees and
older policy bundles remain readable. Start a new campaign for checkpoints
created with the older runtime configuration. See the
[release notes](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/releases/0.2.0a4.md)
and [prefix guide](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/replay-prefix.md).

## Third-party policy and replay packages

Version **0.2.0a2** adds portable JSON bundles and GitHub-hosted package sharing.
Community developers can distribute policy versions and recorded discovery trees;
recipients import them into local SQLite memory with their own agent, quality metric,
task-family contract and acceptance rules. These replay datasets contain observed
search transitions, not model weights. Saved policies can avoid repeated development
when they meet the recipient's quality floor; new tasks still run the recipient's agent.

```bash
dreamrsi list
dreamrsi install OWNER/REPOSITORY
dreamrsi save my-policy-pack --all
dreamrsi publish .dreamrsi/packages/exports/my-policy-pack.dreamrsi.json
```

Use a real package reference for `OWNER/REPOSITORY`. Export requires previously saved
memory; publication requires GitHub CLI and creates a public repository. See the
[complete walkthrough](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/releases/0.2.0a2.md)
for the diagram, executable integration example, and compatibility requirements.

## Basic agent integration

```python
from dreamrsi import Budget, DreamRSI

rsi = DreamRSI(
    agent=lambda task: task.upper(),
    evaluator=lambda answer: float(len(answer)),
    budget=Budget(model_calls=4),
)
result = rsi.run_sync("hello")
print(result.best)  # HELLO
```

This example demonstrates integration, not quality improvement. Stateful adapters support
generate/evaluate/refine tasks. Replay uses recorded outcomes without new discovery calls.
`LLMPolicyDeveloper` writes and iteratively rewrites executable source from measured feedback.
Version 0.2.0a1 adds opt-in persistent champion reuse and more conservative replay
cost accounting. It can replay already recorded runs without a new training call,
delays holdout collection until a replay-improving candidate exists, and stops
repeated policy revisions. The default promotion gate requires paired raw-quality
and probe evidence; applications needing score-only decisions can explicitly use
`ReplayOnlyGate`. An end-to-end cost advantage for an LLM discovery agent has
not been proven.
The SDK includes configurable Docker-free policy interpreters, optional process execution,
held-out validation, SQLite recovery and reported token/USD accounting.

## Two measured experiments

In the [Bonsai Q2 experiment](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/experiments/ternary-bonsai-diverse-2026-09-24.md),
a real local model wrote executable policy code while a deterministic agent solved
the tasks. Across 64 fresh fixture tasks, counted operations fell from 256 to 248
with raw quality 0.9 throughout. This is a logical-operation proxy, not a measured
token or dollar saving for an LLM agent.

In the [GPT-6 Luna xhigh experiment](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/experiments/luna-xhigh-discovery-v1.md),
the real model both solved tasks and developed policies. Deployment fell from
24 to six model requests across six held-out tasks, but 94 preparation requests
made the full path **100 versus 24**. Mean reported score was slightly lower.
This demonstrates learning and reuse, not an all-in economic win.

- [Full documentation and roadmap](https://github.com/TheAstrayDev/dream-rsi-sdk#readme)
- [Bonsai Q2 experiment](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/experiments/ternary-bonsai-diverse-2026-09-24.md)
- [GPT-6 Luna experiment](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/experiments/luna-xhigh-discovery-v1.md)
- [Integration and recovery](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/research-loop.md)
- [Sandbox configuration](https://github.com/TheAstrayDev/dream-rsi-sdk/blob/main/docs/sandbox.md)
- [Issues](https://github.com/TheAstrayDev/dream-rsi-sdk/issues)
- [Original Dream-RSI paper](https://arxiv.org/abs/2609.14858)

Alpha APIs may change. Generated source runs in a bounded Python-syntax subset, not
arbitrary Python. External state isolation and remote cancellation require adapter support.
Usage caps depend on accurate provider reports and ceilings. Research-scale performance
and broad generalization remain unverified.
