Metadata-Version: 2.5
Name: labloop
Version: 1.0.3
Summary: Keep a change only if it measurably helps: an experiment loop for agent-driven research.
Project-URL: Homepage, https://github.com/plicara/labloop
Project-URL: Repository, https://github.com/plicara/labloop
Project-URL: Issues, https://github.com/plicara/labloop/issues
Project-URL: Changelog, https://github.com/plicara/labloop/blob/main/CHANGELOG.md
Project-URL: Roadmap, https://github.com/plicara/labloop/blob/main/ROADMAP.md
Author: Plicara Labs
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: ablation,agents,experiments,machine-learning,research
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Testing
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: mypy>=1.10; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Description-Content-Type: text/markdown

# labloop

[![CI](https://github.com/plicara/labloop/actions/workflows/ci.yml/badge.svg)](https://github.com/plicara/labloop/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/labloop)](https://pypi.org/project/labloop/)
[![Python 3.10–3.14](https://img.shields.io/pypi/pyversions/labloop)](https://pypi.org/project/labloop/)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-green)](https://github.com/plicara/labloop/blob/main/LICENSE)

**Keep a change only if it measurably helps.**

An experiment loop for agent-driven research. Point it at a command that runs
your experiment and a command that changes your code, and it will run trials
under a wall-clock budget — keeping the changes that improve your metric and
reverting everything else.

Every trial is recorded, including the failures. `git log` only remembers what
was kept, and the reverted attempts are most of the information.

```bash
pip install labloop
```

The default proposer sandbox requires Linux, Bubblewrap, and working unprivileged user namespaces. On macOS or Windows, `run` requires an explicit `--sandbox none` opt-out; use it only with code you trust. See [sandbox setup and limitations](#sandboxing-the-proposer). Hosted-model proposers also need `--sandbox-network`.

## Use

In a fresh repository, `labloop init` gitignores the ledger, writes a
stand-in experiment if you have none, and prints the exact first commands.
Then: check the experiment gives the same answer twice, take a baseline, then let an
agent iterate against it:

```bash
labloop noise --run "python train.py" --metric val_loss

labloop baseline --run "python train.py" --metric val_loss

labloop run \
  --run "python train.py" \
  --metric val_loss \
  --propose "my-agent --edit train.py" \
  --budget 300 \
  --trials 50
```

```
[+] trial   0        2.431    41.2s          (baseline)
[+] trial   1        2.298    38.9s  a1f4c02
[-] trial   2        2.355    39.4s
[T] trial   3           --   300.0s
[+] trial   4        2.201    40.1s  7bd9e13

best val_loss: 2.201 (trial 4)
```

`+` kept, `-` reverted, `T` timed out, `!` crashed, `?` no metric found,
`~` the metric was `nan` or `inf`, `=` the proposal changed nothing,
`H` the proposal or measurement changed the harness or ledger, `^` interrupted.

The first step is not ceremony. Keep-or-revert is only as good as the metric
holding still, and [most of what can go wrong](#check-your-metric-holds-still)
starts there.

Or from Python:

```python
from labloop import Experiment, Goal, Loop

exp = Experiment(
    run="python train.py",
    metric="val_loss",
    goal=Goal.MINIMIZE,
    budget_seconds=300,
    propose="my-agent --edit train.py",
    protect=("eval.py", "data/holdout"),
)

loop = Loop(exp)
loop.baseline()
loop.run(trials=50)
```

## How it decides

Each trial runs your `propose` command, then your `run` command, then reads the
metric from the output. The change is committed only if the metric beat the
incumbent. Anything else is discarded:

| Outcome | Meaning |
| --- | --- |
| `kept` | Metric improved. Committed. |
| `reverted` | Metric was worse, or tied. |
| `failed` | The command exited non-zero. |
| `timed_out` | Exceeded its budget. Process group killed. |
| `no_metric` | Ran clean but printed no metric. |
| `not_finite` | The metric was `nan` or `inf`. Nothing compares to it. |
| `no_change` | The proposal edited nothing, so there was nothing to measure. |
| `harness_changed` | The proposal or measurement changed protected files or the ledger; its metric was discarded. |
| `interrupted` | Stopped by hand partway through. |

Four details that matter:

- **A tie is not an improvement.** Equal scores revert, so the loop never
  accumulates neutral churn.
- **A missing metric is not a bad score.** A broken experiment and a poor
  result are different events and are recorded differently. So are a crash, a
  diverged run that printed `nan`, and a proposal that edited nothing — each
  sends you somewhere different, so each gets its own outcome.
- **The loop refuses to start on a dirty tree.** It reverts by discarding, so
  uncommitted work would be destroyed.
- **A metric from a changed harness is not a result.** See below.

`--budget` is how long the experiment may run. An agent that thinks for longer than the experiment takes is ordinary, so give the proposal its own with `--propose-budget SECONDS` rather than raising both. Either one that overruns is killed with its POSIX process group and recorded as `timed_out`. An unconfined child can escape that group by creating a new session; the proposer sandbox's PID namespace closes that route on Linux. Measurement commands run outside that namespace. Windows termination is best-effort and is not covered by the Linux CI matrix.

Only one trial executes against a ledger at a time. By default, `run` holds the lock for its entire run and refuses a competing writer with the holder's pid. With `--wait`, it waits for access and releases the lock between trials, so waiting runs can interleave complete trials, with no fairness guarantee. The lock dies with its process. This is a same-host advisory lock: do not run simultaneous writers on different machines against a shared ledger.

It also stops when it stops learning. Ten trials in a row that produce no metric
at all — a mistyped `propose` command, an agent that never applies its edit —
end the run rather than spend the rest of an overnight budget failing
identically. Occasional failures don't count; only an unbroken run of them does.
Change it with `--give-up-after N`, or `0` to run regardless.

## Research directions

Autoresearch grows a single thread of commits; its author has said the next
step is many. A direction is a parallel line of inquiry over the same shared
ledger, with its own incumbent:

```bash
labloop branch wide-lr --from-trial 7
git worktree add ../wide-lr -b labloop/wide-lr <trial-7-commit>
cd ../wide-lr && labloop run --direction wide-lr --ledger <shared> ...
```

The fork starts from the kept trial's metric — its first attempt has to beat
where it forked from, not zero, and not the parent's later progress. Trials
carry their direction, indices stay globally unique, and `labloop log`
reports each direction's best side by side. A proposer's brief contains only
its own direction's history.

Directions using `--wait` share a per-trial ledger lock: their commands do not execute simultaneously, but waiting loops in separate worktrees can alternate complete trials. A default, non-waiting run holds the lock until its full run finishes. Read-only queries need no lock. Do not run simultaneous directions in the same worktree.

## Crashes and resuming

Every run records the spec it started under — command, metric, goal,
budgets, protected files — as a manifest line in the ledger (never the
environment, which is where credentials live). If the machine dies mid-run:

```bash
labloop resume --trials 20
```

continues under the recorded spec, same incumbent, same numbering — nothing
retyped, nothing drifted. And because the metric name and goal define what
the recorded numbers mean, a later run that changes either is refused with
the field named: comparing a `val_loss` being minimized against an
`accuracy` being maximized would mix measurements and tell no one.

## Reading the metric

Two formats, no configuration. The last `key=value` or `key: value` occurrence wins. If there are no such pairs, the last matching JSON-line value wins. Use one format consistently: a later JSON line does not override an earlier key/value pair.

```
val_loss = 1.234        # key=value or key: value
{"step": 40, "val_loss": 1.234}    # a JSON object on its own line
```

## Check your metric holds still

Keep-or-revert assumes that a change in the metric means a change in the code.
If your experiment scores differently run to run, that assumption is false, and
the loop will commit the luckier draws and report them as progress.

Find out before you start:

```bash
labloop noise --run "python train.py" --metric val_loss --repeat 6
```

```
val_loss: 0.857473 to 1.11126 over 6 identical runs
spread: 0.253788   standard deviation: 0.0983

An improvement smaller than 0.253788 is a difference this experiment has already
produced without any change to the code, so the loop would be selecting lucky
runs. Best is to remove the variance — fix the seed, average more, hold the data
still. Failing that:

    labloop run --min-delta 0.253788 --confirm ...
```

Nothing changed between those runs. Any "improvement" below the spread is the
loop picking a good roll of the dice. The spread is what to clear, but it widens
with every extra run; the standard deviation is the one to compare against a
later measurement or another experiment.

Four worked experiments in [the cookbook](cookbook/noise-across-recipes.md)
measured 22%, 8.9%, 0.3% and 0% on one machine — including two timing
benchmarks that differ by 70× — so this is not a number to assume.

**Removing the variance is the real fix.** Fix the seed, average over more
data, hold the split still. Two settings help when you can't:

- `--min-delta D` — the metric must improve by more than `D` to count. Attacks
  how *often* a fluke is kept, and costs nothing.
- `--confirm` — re-run before keeping, and keep only if it wins twice. The
  incumbent then advances to the weaker of the two measurements, so a lucky
  draw doesn't set a bar only luck can clear. Attacks how *far* the fluke
  drifts, and costs one extra run per candidate win.

Measured on a metric that is pure noise, where every kept trial is false by
construction — 60 trials, averaged over 400 runs:

| Setting | Improvement claimed | False keeps | Experiment runs |
| --- | --- | --- | --- |
| default | 23.5% | 4.8 | 60 |
| `--min-delta` (1 sd) | 20.7% | 2.4 | 60 |
| `--confirm` | 12.9% | 4.8 | 72.7 |
| both | **9.9%** | **2.1** | 66 |

They work on different halves of the problem, and are cheaper together than
`--confirm` alone — `--min-delta` rejects most candidates before they earn a
second run. Neither makes a noisy metric safe. They make it less wrong.

## Protecting the measurement

A keep-or-revert loop rewards whatever moves the metric, and your `propose`
command can reach the evaluator. Agents take that route: published runs have
seen them overwrite test cases and memorize evaluation answers rather than
improve anything.

Name the files that define the measurement and labloop digests them with SHA-256 before the proposal, after the proposal, and after each completed measurement, including confirmation runs. Baselines and noise calibration also check their protected files before and after measurement:

```bash
labloop run \
  --run "python eval.py" \
  --metric val_err \
  --protect eval.py \
  --protect data/holdout \
  --propose "my-agent --edit train.py"
```

```
[+] trial   0             1     0.0s  (baseline)
[H] trial   1            --     0.0s  (proposal modified the harness: eval.py)
[+] trial   2        0.3333     0.0s  7c599cd
```

A pattern may name a file, a glob, or a directory — a directory covers the
whole subtree, which is usually what frozen evaluation data needs. Renames,
deletions, and added files all move the digest, because memorizing answers
means adding files and not only editing them.

**Digests detect changes to the watched files; they do not prove measurement validity.** A changed protected set produces `harness_changed` instead of a score, even if the command also failed or timed out. The note identifies the phase: `proposal`, `run`, `confirmation run`, or `baseline run`. The ledger is checked at the same points without needing `--protect`; detected rewrites, deletions, and leaf symlink replacements are restored from trusted pre-command bytes before recording the rejection. If its directory was redirected or the ledger was replaced by a directory, recovery stops for manual inspection.

Baselines reject a changed harness without creating a new incumbent. A baseline that started with a clean Git tree restores that tree; one that started dirty preserves uncommitted work and asks you to inspect it. Noise calibration holds the ledger lock across all repeats, aborts on a protected-file or ledger change, and returns no statistics or trial records. It restores an altered ledger (or removes one created by the command), but leaves the worktree for inspection.

Matching digests establish only that files matched at the check points. They cannot detect a command that changes a file, uses it, and restores it before exiting. They also cannot authenticate stdout metrics, prevent monkey-patching, or cover unwatched dependencies. Interrupted commands do not reach these post-command checks. The [evaluation-isolation proposal](docs/evaluation-isolation.md) describes a stronger boundary; it is not implemented.

If the incumbent in your ledger was measured under a different digest, the loop
stops rather than compare two numbers that came from different measurements.

If the complete protect set matches no files, startup fails. Individual unmatched patterns alongside valid matches are not currently rejected, so check each path. When something moves, the trial names the file: `proposal modified the harness: data/holdout.csv` tells you where to look.

**Protect the measurement, not the directory it lives in.** If your evaluator writes a cache or a log inside a protected path, the current measurement is rejected. Caches are artifacts; keep them somewhere you are not protecting.

## Sandboxing the proposer

Only the `--propose` command is sandboxed. The `--run` measurement command executes on the host and can load code the proposer just changed. For adversarial or unknown code, run the entire experiment on a disposable, credential-free machine or VM; this proposer sandbox is not end-to-end containment.

On Linux, labloop runs the propose command inside **bubblewrap**: it can read the whole machine, write the worktree and explicitly granted evidence directories, and gets its own process and network namespaces. The project's `.git` is bound read-only and known Docker, Podman, and containerd socket paths are masked. Custom socket paths are not covered. This is the default; `--sandbox none` opts out.

The wrapper explicitly drops all Linux capabilities (`--cap-drop ALL`) and starts a new session (`--new-session`) to detach from the controlling terminal. These harden the proposer boundary; they do not sandbox the measurement command.

It needs bubblewrap and unprivileged user namespaces:

```bash
sudo apt install bubblewrap
```

Some Ubuntu installations restrict unprivileged user namespaces through AppArmor. Follow your host's administrator-approved policy; disabling `kernel.apparmor_restrict_unprivileged_userns` is a machine-wide security change, not an application-local setting. CI changes it only on disposable runners.

Without them the loop **refuses to start**. A startup self-check verifies a worktree write, a denied outside write, and each declared evidence write before trials begin. It is a capability check, not a comprehensive security proof. The Python API is confined by default too; `LABLOOP_SANDBOX=none` opts out.

**The network is off by default.** A proposer that calls a hosted model needs
`--sandbox-network`; a local or scripted one does not. Off means a compromised
proposer cannot send what it reads over IP — but Unix-domain sockets stay
reachable, so it is "no IP egress", not "cannot talk to anything".

A sandboxed proposer's durable evidence belongs under the fixed root `~/.local/state/labloop/evidence`. `--sandbox-write PATH` (repeatable) grants an existing strict descendant of that root, outside the worktree and not an ancestor of it. The root itself is refused. Symlinks cannot redirect the root or grant directories outside it, and both the evidence path and worktree are canonicalized before checking.

```bash
mkdir -p "$HOME/.local/state/labloop/evidence/my-experiment"
labloop run --run "python train.py" --metric val_loss \
  --propose "my-agent --edit train.py" --sandbox-network \
  --sandbox-write "$HOME/.local/state/labloop/evidence/my-experiment"
```

Tell the proposer where to write its transcripts; the flag grants access but does not create or capture logs. Choose a separate directory per experiment and treat its contents as untrusted proposer output. Agent caches must be configured inside the worktree or a permitted evidence directory; arbitrary writes to `~/.config` or credential directories are not supported. Existing manifests that grant other locations must be replaced by a new `run` invocation with an allowed path before using `resume`.

Two things it does not do: reads are open by design (keep secrets off the
machine), and bubblewrap shares the host kernel — it is not a virtual machine.

## What the proposer is told

A proposal command that gets no feedback is guessing. Before each attempt
labloop writes the trial history to a JSON file and puts its path in
`$LABLOOP_BRIEF`:

```json
{
  "trial": 5,
  "metric": "val_loss",
  "goal": "minimize",
  "incumbent": 1.5,
  "protected": ["eval.py"],
  "counts": { "kept": 2, "reverted": 2, "failed": 1 },
  "history": [
    {
      "index": 1, "outcome": "reverted", "metric": 2.0,
      "why": "reverted: val_loss 2 tied the incumbent, and a tie is not an improvement"
    },
    {
      "index": 3, "outcome": "reverted", "metric": 3.0,
      "why": "reverted: val_loss 3 did not beat 1.5; lower is better"
    }
  ]
}
```

The `why` is the part the proposer can't work out for itself. `reverted` is a
label; *tied the incumbent, and a tie is not an improvement* is something to
act on. Failures carry the tail of their output, so an agent can see the stack
trace that killed its last three attempts.

For a one-line proposal command that doesn't want to parse JSON, the same
essentials are in `$LABLOOP_METRIC`, `$LABLOOP_GOAL`, `$LABLOOP_INCUMBENT`
(empty when there is nothing to beat yet) and `$LABLOOP_TRIAL`.

The brief is written by labloop and read by the proposal, never the reverse.
Agents handed a memory file they can write have been seen leaving notes for
their future selves, which turns persistent memory into a way around the
harness rather than a record of it. The agent learns what happened without
getting to decide what happened.

Pass `--no-brief` to turn it off. The file is written outside the working tree
either way, so it never dirties the tree or lands in a commit.

## What gets committed

A kept trial commits the paths changed by the proposal and `labloop-history.jsonl` — a sparse decision log with one compact line per trial, reverted ones included. Paths are captured before measurement, including individual files in newly created directories. Their contents are committed after measurement: an evaluator that modifies a proposed file changes what gets committed. Keep evaluator artifacts in separate paths.

Nothing else. Whatever the run produced beyond the proposed change —
checkpoints, logs, caches — is discarded after the trial is judged, exactly as
it already was on a reverted trial. A training run that writes a checkpoint
per trial would otherwise turn an overnight loop into a repository of hundreds
of gigabytes, and the commit would stop meaning "the change that improved the
metric".

If you want an artifact to survive across trials (a download cache, a
warm-start checkpoint), put it in `.gitignore`: ignored files are never swept
and never committed. If you gitignore the decision log itself, it is still
written locally but left out of commits — a stated preference is respected.

## The ledger

Trials append to `labloop.jsonl` — one JSON object per line, readable while the
run is still going. Query it without leaving the tool:

```bash
labloop log --metric val_loss              # replay with per-direction bests
labloop log --json                         # one strict JSON object per trial
labloop log --outcome reverted --json      # only what was thrown away
labloop log --since-trial 40 --direction wide-lr
labloop log --compare main wide-lr         # refuses if their harnesses differ
```

```python
from labloop import Goal, Ledger

ledger = Ledger("labloop.jsonl")
ledger.summary()        # {'kept': 7, 'reverted': 31, 'timed_out': 2, ...}
ledger.best(Goal.MINIMIZE)
```

## Prior art

The keep-or-revert loop is the idea behind
[Karpathy's autoresearch](https://github.com/karpathy/autoresearch), which
wires it directly into single-GPU nanochat training. `labloop` is not that
project and is not affiliated with it. It generalizes the loop: any command,
any metric, no GPU assumption, with the trial history as a queryable artifact
rather than scrollback.

## Status

Released as 1.x. The Python package uses only the standard library; the default proposer isolation additionally requires Linux and Bubblewrap. See [the audit notes](docs/audit-2026-09-14.md) for verification scope and known limitations.

## License

Apache-2.0.
