Without flip
A polished answer with no trail
A four-page report says public EV-charger reliability improved. Twelve URLs sit at the bottom. Six weeks later, nobody can tell what was measured, what was inferred, or what was never checked.
Reporter's notebooks for agents v0.21 draft spec
Deep-research agents hand you a fluent report and a list of URLs. flip keeps the record underneath: every source captured and judged, every claim at an honest status, every open question named, in plain files that outlive the session.
Without flip
A four-page report says public EV-charger reliability improved. Twelve URLs sit at the bottom. Six weeks later, nobody can tell what was measured, what was inferred, or what was never checked.
With flip
The narrow claim is verified against two captured sources. The national trend stays unresolved, with the surfaces checked and the evidence that should reopen the question recorded.
What would you use it for?
flip shapes the record to the work. The agent chooses the matching notebook kind; you keep speaking in questions and outcomes.
| What the work needs | Notebook kind |
|---|---|
| Screen whether an angle is worth pursuing | scout |
| Run one question to ground | pursuit |
| Survey a field or prepare a publishable review | research-review or lit-review |
| Recompute, reconcile, or investigate data | data-investigation |
| Prepare an evidence-backed choice | decision-packet |
| Record forecasts and what would settle them | forward-set |
| Maintain a shared source record | ledger |
| Work inside confidential boundaries | engagement |
See it happen
The flipbook replays a completed investigation. These three moments—and every command underneath them—are regenerated with the real CLI when the site is built.
You ask
Track where OKF came from, its precedents and gaps, and whether there is evidence of a roadmap.
The evidence gate refuses
cannot verify: the corroboration bar is not met
The next agent inherits
The origin is bounded, the gap map is corroborated, and the roadmap finding says exactly where it does and does not reach.
The gap
Deep-research systems can return fluent, cited reports while leaving important branches unsearched, presenting one side of a question, or attaching citations that only partly support the prose. Six weeks later, a URL list and a polished answer cannot tell the next agent what was actually checked, what remained weak, or which follow-on question the evidence opened.
Bracketed ids open this site's own flip notebook, where each claim on the page is recorded with its sources and status.
On 100 open-web research tasks, the best system in one independent benchmark got about half the job done—0.55 F1, covering only around half the searches the tasks required (LiveDRBench) [C8]. In another, deep-research systems argued one side of debate questions 55–95% of the time, and their citations fully supported the prose only 31–79% of the time (DeepTRACE, August 2025 snapshot) [C9].
flip makes the route taken—and the routes still owed—durable, inspectable, and continuable.
What changes
01
Open questions, failed probes, and conditions for revisiting an answer survive the session, so the next run continues instead of starting over.
02
Sources are captured and judged as separate acts. A lead can sit next to a verified claim without pretending to be one.
03
Stable ids, named contributors, and a preserved history let another agent or person continue without reconstructing the work from chat—or silently overwriting how it changed.
Observed use
An agent captured four NJ enrollment workbooks, found file oddities, recomputed totals two ways, and opened the more consequential follow-on question the popular framing missed.
NJ schools notebook →A literature review froze criteria before searching and kept the full denominator: 2,600+ identified, 31 examined, seven advanced, four excluded, and three included—including a canonical paper excluded on license alone.
Literature review →An EV-charging pursuit distinguished failed visits from measured uptime, answered the narrower question, kept the national trend unresolved, and recorded exactly what evidence should reopen it.
EV charger pursuit →Forty-three unattended sessions over nearly 67 hours retained 380 sources, 82 claims, 47 questions, corrections, bounded nulls, and two deliberately unconfirmed load-bearing claims.
Autonomous-loop case study →What is actually enforced
Every fragment below is real output, generated by running the CLI when this page was built.
At rest
A flip notebook is an Open Knowledge Format v0.2 knowledge bundle at rest—not an export target [C6]. Every source, claim, decision, question, and session is one Markdown page with YAML frontmatter and an immutable id. Raw captures and hashes establish custody; events are append-only JSONL; generated views provide the hot state.
Open the same directory in Git, Obsidian, any Markdown editor, or an OKF consumer. Export it as BagIt, CSL JSON, render JSON, or a policy-filtered public OKF copy. Public export can withhold private source custody and event history; the workflow still owns source rights and licensing decisions.
Where it fits
flip is the durable research record beneath retrieval, orchestration, and publishing tools. These differ in kind, not in quality.
| Dimension | flip notebook | Plain markdown in a repo | PKM vault | RAG / vector store |
|---|---|---|---|---|
| Canonical at rest | Markdown pages + append-only JSONL | Markdown files | Files plus application state | Embeddings in an index |
| What it enforces | Custody, separate judgment, verification gates, and recorded tests | Nothing—convention only | Schema and citation affordances | Retrieval relevance, not evidence lineage |
| What it is for | A continuable research record | Flexible authored documentation | Human knowledge work | Finding likely-relevant material |
| If the tool disappears | A valid OKF bundle and readable directory remain | The files remain | Files remain; derived app state may not | The index goes with the tool |
What the evidence supports
The strongest objection is that this is ceremony. Source grading and claim metadata can raise the cost of every turn without improving what an agent concludes. A careless model can fill fields carelessly.
The answer is structural, not magical. Grading is a separate recorded act by a
named actor. Ungraded sources count toward nothing. flip doctor
exposes gaps and inconsistencies. That makes carelessness inspectable; it does
not prevent it.
The public examples show real notebook behavior, correction, long-running work, and honest unresolved states. No controlled test yet shows that flip-backed agents produce better conclusions than agents without notebooks. The measured claim today is narrower: on one 507-page notebook, cold orientation through generated views cost about 40K tokens, motivating bounded frontier views—not a blanket claim that a CLI always uses fewer tokens.
flip's provenance vocabulary has not been submitted to, reviewed by, or endorsed by OKF maintainers. The spec is draft v0.21, not 1.0 or frozen. Migration exists because the format has moved and may move again.
Where this stands
Built and maintained by Marc Lavallee. MIT licensed; the format and software are public, and issues are welcome.
This site's own lineage
Generated from this site's flip notebook by
flip export json at build time. Unconfirmed and superseded claims
stay visible rather than disappearing from the story.