Metadata-Version: 2.4
Name: betaloop
Version: 0.9.1
Summary: A reusable, storage-free ReAct agent kernel (engine + framework + protocols) with optional capability bundles.
License: MIT License
        
        Copyright (c) 2026 betaloop contributors
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Keywords: agent,llm,react,kernel,tool-calling,mcp
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Application Frameworks
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx>=0.24
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
Requires-Dist: ruff>=0.4; extra == "dev"
Dynamic: license-file

# betaloop

A reusable, **storage-free** ReAct agent kernel: the engine, framework and
protocols that drive a tool-calling agent, plus optional capability bundles. It
knows nothing about how runs are stored (or even whether they are) — persistence
is an optional `EventSink` a host plugs in. Any application (a thesis-writing
platform, a coding agent, ...) implements its own tools + prompt + storage and
reuses this kernel.

## Core (zero I/O, zero business)
- `runtime` — `AgentRuntime` ReAct loop + event stream. Supports cancellation
  (`run(..., stop=Event|callable)` → `cancelled` event, `status="cancelled"`;
  checked between streaming deltas too, closing the in-flight model stream
  instead of paying for a response nobody wants) and streaming
  (`LLMConfig(stream=True)` → `assistant_delta` events while the
  model generates; a final full `assistant` event always follows). Every
  `tool_call` event is announced before any of the step's tools execute, so a
  slow tool never hides what is pending on the frontend; result events still
  follow in model order. Every model call emits a
  `usage` event — prompt/completion/total tokens, that call's
  cost, and context fullness (`context_tokens`, `context_chars`,
  `context_window`, `context_percent` when `LLMConfig(context_window=...)` is
  set) — so a frontend can show live token/context gauges; `RunStats` and the
  host `done` event carry the cumulative breakdown. Run budgets
  (`max_cost` / `max_total_tokens`) cut a runaway run short with
  `status="budget_exceeded"` — no further tool execution, no further model
  calls — and a host may seed `stats` with prior-conversation totals to
  budget across runs. A stuck model reissuing the *identical* (tool, args)
  call more than `repeat_call_limit` (default 3) times gets an inline nudge
  inside that tool's result, feeding self-correction (`None` disables).
  Malformed tool-call arguments come back to the model as failed tool
  results instead of executing with empty/wrong args. A run that exhausts
  its step budget gets a forced toolless wrap-up call (`tool_choice="none"`)
  so it ends with the model's summary, reporting `status="max_steps"` (and
  never executing the stubborn model's further tool calls).
  `LLMConfig(temperature=..., max_tokens=...)` are forwarded on every call,
  and a generation cut off by the token cap (`finish_reason="length"`) is
  marked — the notice rides in the final text, the per-call `usage` event
  carries `finish_reason`, and a tool_call whose arguments JSON was truncated
  gets a "cut off by max_tokens, re-issue the call" error instead of a bare
  "invalid JSON" that invites a byte-identical retry. An empty response (no
  text, no tool calls) is retried once and otherwise ends the run with
  `status="empty_response"` instead of an empty "success". A missing
  `tool_call` id is synthesized at one point (streamed and non-streamed
  alike), so a gateway that omits ids never poisons the next request with
  `tool_call_id: null` — and streamed calls no longer mint colliding
  `call_0`-style ids across steps. Synthetic run endings (max-step /
  budget notices) are recorded like any assistant turn, so sinks and the
  next run's replay see why the run stopped. Tool calls execute in model
  order with consecutive READ tools parallel and every WRITE/META tool alone
  (no write races); `ToolSpec(timeout=...)` cancels a hung call. A mid-run
  context budget (`context_budget`, default 400k chars) shrinks old tool
  results head+tail so long runs don't blow the context window. Sinks are
  error-isolated (a broken display/record sink logs instead of killing the
  run; `strict_records=True` opts record failures back into fatal).
  `AgentRuntime(envelope=True)` emits the `run_start`/`done` envelope itself
  for hosts driving the runtime directly, and `AgentRuntime(http_client=...)`
  (forwarded by `AgentHost`) reuses one host-owned `httpx.AsyncClient` across
  runs — connection pooling, limits, proxy/verify config for service hosts.
- `llm` — OpenAI-compatible chat client (retry / jittered backoff / response-shape
  validation / `Retry-After`-aware 429 handling / fatal-4xx fail-fast) + SSE
  streaming helpers; transports normalize usage to
  `prompt_tokens`/`completion_tokens`/`total_tokens` across chat-completions
  and Responses shapes
- `tools` — `ToolRegistry` (register / unregister / dispatch / mode filtering /
  argument validation / per-tool timeout) + pre-dispatch **middleware** via
  ``add_middleware`` (audit / quota / human-in-the-loop confirmation of write
  tools)
- `actions` — `Action` + `UndoEngine` (pure, storage-free undo; reverters may
  be sync or async)
- `memory` — `replay_messages` / `recap_text` / `window_with_recap` /
  `run_timeline` (reconstruct a stored turn's ordered event timeline) +
  `MemoryProvider`; replay reconciliation is two-sided *and* window-safe
  (a tool row whose calling assistant fell outside the window is dropped,
  not sent as an orphan first message)
- `events` — `EventSink` (display + record channels) + SSE serialization
  (`to_sse` degrades non-JSON values via `str()` — a `datetime` inside a
  tool's `ui` payload can't crash the host's SSE layer)
- `context` / `modes` — `AgentContext` + `AgentMode` / `ToolCategory`;
  host-defined modes via ``register_mode(name, categories)`` (unknown modes
  raise instead of silently degrading to read-only). `AgentContext.shared`
  is per-run state shared *by reference* with subagent contexts — the
  vehicle for cross-context coordination (the workspace stale-file guard's
  revision map, the run's cancellation handle)

## Optional bundles (`betaloop.bundles`)
- `host` — `AgentHost` host-adapter framework: message assembly, run envelope
  (`run_start`/`done`), error funneling, `StoreSink` (persist via a store),
  `undo_run` (zero-config: bundled tools register reverters keyed by their
  action kinds, and reverters see the host's `extra` context; a reversion
  whose *status mark* fails surfaces in the report instead of being
  swallowed), `DictToolAdapter` (wrap a dict-based tool system — specs
  without a handler log a warning instead of silently vanishing, and a
  `reverters=` map wires custom action kinds into undo).
  Eliminates the per-host boilerplate round 1
  left behind. `host.run(..., stop=...)` forwards cancellation to the runtime;
  `AgentHost(max_cost=..., max_total_tokens=..., repeat_call_limit=...)`
  forwards the runtime's budget / repeat guards to every run it builds;
  `AgentHost(http_client=...)` shares one HTTP client across runs, and
  `AgentHost(capture_actions=False)` turns off `StoreSink`'s automatic
  action capture for hosts that persist actions themselves.
- `store` — `RunStore`/`ConversationStore`/`BlobStore` Protocols +
  `JsonlRunStore` (default, **zero-database** JSONL + content-addressed blob
  spillover). Hosts wanting a DB implement the Protocols; the default needs
  none. Action values larger than `spill_threshold` (default 8KB) externalize
  to blobs and rehydrate transparently on read; id counters are in-memory so
  appends don't rescan the stream. `list_actions(subagent=...)` filters by
  subagent during the fold (unselected rows skip blob rehydration — the
  subagent engine's snapshots stay O(one worker) on long runs); torn lines
  log a warning and count in `dropped_lines` instead of vanishing silently;
  `fsync=True` flushes each append for crash-durability.
- `subagents` — `SubagentEngine` + `SubagentRoster`/`SubagentSpec` + `delegate`
  tool: isolated worker agents the orchestrator hands subtasks to, tagged so undo
  still reverts them while the orchestrator's context stays lean.
  `delegate_parallel` fans independent tasks out concurrently (bounded by
  `max_parallel`, failures isolated per agent, duplicate agents in one batch
  rejected — they would claim each other's actions); `SubagentEngine(on_subagent_event=...)`
  streams live `subagent_progress` heartbeats to a host push channel (text,
  args and summaries capped — a 50KB `write_file` payload never rides the
  callback); `SubagentSpec(transport=...)` routes a subagent to a different
  endpoint. `make_delegate_tool(engine, timeout=...)` caps one delegation's
  wall time so a hung worker cannot hold the orchestrator's step forever.
  Budget caps apply per subagent run (each delegation gets its own
  `max_cost` / `max_total_tokens`) while each delegation's spend folds into
  the parent run's `done` event, stats and store row (with a
  `subagent_*` breakdown); cancellation of the orchestrating run
  propagates into an in-flight subagent (its stop handle rides in
  `ctx.shared`); a delegation's own mutations are identified by an
  id-membership snapshot, so stores with opaque (non-integer) action ids
  count them correctly.
- `admin` — `tool_categories` / `list_tools_admin` / `list_tool_packages_admin` /
  `check_packages` over a registry + display packages (admin-panel source;
  packages are grouping only, tools stay per-name togglable).
- `patch` — `apply_patch`: line-oriented multi-file edits via the Codex
  `*** Begin Patch` envelope (add/update/move/delete files, `@@` chunks of
  context/`-`/`+` lines, `*** End of File` anchoring). Chunk location runs a
  four-pass fuzzy ladder (exact → trailing-ws → strip → Unicode-punctuation
  fold); `*** End of File` chunks run the tail-anchored ladder at full
  strength before any forward match, so a whitespace-mismatched tail beats
  an exact look-alike earlier in the file instead of silently editing the
  wrong site; two same-position insertions keep document order. Application
  is all-or-nothing against an in-memory overlay (later
  sections of the same file chain; create-then-edit works), so a bad hunk
  leaves the workspace untouched. Same `file_change` events + one new undo
  kind (`file_delete`).
- `workspace` — sandboxed file I/O + read/write/edit/list/search/glob tools +
  file undo reverters. Directory walks (`list_files` / `search_files` /
  `glob_files`) never follow symlinks — code executed by `run_code` could
  otherwise plant a link to a host file and read it back through a walk,
  bypassing the path guard (symlinks list as an opaque `symlink` type;
  `list_files(dirs=...)` routes the requested folders through the same path
  guard). `read_file` prefixes every line with its 1-based
  number and supports `offset`/`limit` line-window reads. `edit_file` refuses
  ambiguous `old_text` (multi-match) unless `replace_all` is set, returns a
  diff, and falls back to a whole-line fuzzy match (shared ladder,
  trailing-blank aligned like `apply_patch`) when the
  exact substring misses. All write tools guard against stale content: a
  file read this run and changed out-of-band is refused with "re-read it"
  instead of being clobbered; the revision map lives in `ctx.shared`, so the
  guard spans the orchestrator and every (including parallel) subagent of
  the run. `search_files` greps content
  by regex (dir / glob filters) with a 30s tool timeout, a 10s scan budget
  and a 10k-char line cap (catastrophic-backtracking patterns and huge
  minified lines can't hang the loop); `glob_files` matches paths by pattern.
- `todos` — a per-scope task list the agent plans against: `TodoStore`
  (pure container) + `JsonTodoStore` (atomic tmp+rename JSON persistence;
  malformed items are dropped on load instead of crashing the system
  prompt) + `update_todos`/`list_todos` tools (`todo_replace` reverter) +
  `todos_block` for splicing the list into a system prompt.
- `images` — tool-tier image perception, no kernel changes. `image_info`:
  stdlib-only header probe (PNG/JPEG/GIF/BMP/WEBP — dimensions, dpi, color
  mode) answering deterministic questions with zero model calls; it stats
  first and reads only a bounded header window off the event loop.
  `analyze_image`: ONE vision-model call (OpenAI `image_url` data-URL block +
  the question), registered only when a `LLMConfig` is passed — the host's
  main config reuses the main model, a dedicated one routes vision elsewhere.
  Size is checked by stat before reading (a 2GB upload is refused, not
  loaded), file reads run off the event loop, and `detail="auto"` shares a
  cache key with an omitted detail (the API treats them identically — no
  double billing). Images never enter the main conversation (answers are
  memoized per file hash + question), so context budget / trimming / replay
  stay untouched.
- `sandbox` — Python code execution (bubblewrap or passthrough backend);
  output truncation keeps head+tail so tracebacks at the end stay visible.
  `run_code`/`run_file` are WRITE-classified (executing model-written code
  can mutate the workspace): invisible in read-only modes and never run in
  parallel with other tool calls. Their side effects produce no undo
  records (the tool descriptions say so) — durable edits belong in
  `write_file`/`edit_file`/`apply_patch`. Note the contract difference:
  `register_code_tools` takes `workspace_for(ctx) -> root path` (a `str`),
  while the workspace/images bundles take `workspace_for(ctx) -> Workspace`.
- `download` — `download_file(url, path?)`: stream one HTTP(S) resource into
  the workspace (default `downloads/`), the controlled ingress the
  network-isolated sandbox deliberately lacks — fetching is a bounded,
  audited tool while executing model code stays offline. SSRF guard: http/
  https only, redirects followed manually with every hop re-validated, and
  all resolved addresses of every hop must be globally routable
  (`ipaddress.is_global` rejects loopback/private/link-local/CGN/reserved/
  multicast — cloud metadata included); checking all A/AAAA records up front
  narrows DNS rebinding to the TTL window. Size cap by `Content-Length`
  pre-check plus streaming cutoff (a lying header gets cut mid-stream and
  the partial `.part` file removed; the file lands atomically via rename).
  Existing targets are refused (the model picks a new name), so nothing
  undoable is mutated. Defaults: 64MB cap, 120s total budget (kernel
  `ToolSpec` timeout as backstop), 5 redirects; WRITE-classified like the
  other workspace-mutating tools.
- `skills` — markdown skill libraries (`SkillLibrary` flat dir; package-aware
  `SkillPackages` with `RemoteSkillSource` registry mirrors — refresh
  failures clean their staging dir, log, and keep the previous cache) +
  `load_skill` tool. Skills do file/network I/O, hence a bundle — `import
  betaloop` stays zero-I/O (deprecated `betaloop.skills` alias kept; it
  warns on import).
- `mcp` — `MCPManager` + `MCPServerConfig` + `parse_servers`: bridge external
  MCP servers into the
  registry — stdio transport (`command=[...]`, e.g. `npx -y @z_ai/mcp-server`)
  or streamable-http (`url=` + `headers=` for auth, `Mcp-Session-Id` handled
  automatically). `tools/list` pagination (`nextCursor`) is followed, so
  paginated servers don't silently lose half their tools. Spawned servers
  get only a safe env allow-list plus the configured `env` (PATH/locale/
  HOME/TMPDIR) — never all of `os.environ` with its secrets — unless
  `inherit_env=True` opts back into the legacy behavior. URLs log as
  scheme://host only (credentials in the query or userinfo never reach a
  log line). Sessions outlive registry rebuilds and lazily self-heal: a dead
  session is restarted on the next tool call (`revive`), and
  `attach`/`ensure` re-register tools on any fresh registry (`ensure` is the
  idempotent variant for hosts that cache registries across runs).
  `tool_allowlist`/`tool_blocklist` filter by the server-side tool name.
  `readOnlyHint` annotations map to the READ category; MCP tools carry no undo
  reverters. Stdlib-only
  (newline-delimited JSON-RPC / plain POST + SSE), secrets stay host-side:

  ```python
  from betaloop.bundles import MCPManager, MCPServerConfig

  manager = MCPManager([MCPServerConfig(
      name="zai", command=["npx", "-y", "@z_ai/mcp-server"],
      env={"Z_AI_API_KEY": "...", "Z_AI_MODE": "ZHIPU"},
      default_category="read")])
  await manager.attach(registry)   # registry gains zai__* tools
  ```

## Install
```bash
pip install betaloop
# local dev (editable + test/lint deps):
pip install -e ".[dev]"
```

## Storage model
The kernel stores nothing. A host provides:
- an **`EventSink`** (write side) — persists message records however it likes
  (DB / file / nowhere);
- a **`MemoryProvider`** (read side, optional) — replays prior turns;
- an **`UndoEngine`** fed from wherever the host kept actions.

## Minimal host sketch
```python
from betaloop import AgentContext, LLMConfig, ToolRegistry
from betaloop.bundles import AgentHost, JsonlRunStore
from betaloop.bundles.workspace import Workspace, register_file_tools

registry = ToolRegistry()
# workspace_for(ctx) -> Workspace (a sandboxed root per user):
register_file_tools(registry, lambda ctx: Workspace(f"/data/{ctx.user_id}"))
store = JsonlRunStore("/var/lib/myapp/agent")                       # zero-DB default
host = AgentHost(registry,
                 LLMConfig(model=..., base_url=..., api_key=...),
                 store, build_system_prompt=my_prompt_builder)

ctx = AgentContext(run_id=rid, user_id=uid)
async for event in host.run(ctx, task, history=prior_turns):
    ...  # forward run_start / step / tool_call / tool_result / done to your frontend
```
A host supplies only its **tools**, **system prompt**, and (optionally) a store
backend — the engine, persistence, run envelope, undo, and (via `subagents`)
delegation are all reused.

## More
- `examples/minimal_host.py` — a runnable, offline minimal host (tools → run →
  events → undo); runs in CI.
- `examples/live_host.py` — the same flow against a real OpenAI-compatible
  endpoint (set `BETA_API_KEY` / `BETA_BASE_URL` / `BETA_MODEL`; no default
  endpoint, it never spends tokens by accident).
- `CHANGELOG.md` — what changed and when.

## License
MIT
