Lanzadera is an on-demand commuter shuttle for Spanish suburbs whose core covenant is a +/-6-minute pickup window.
What this run concluded
Every number on this page is derived from the append-only Journal — nothing was written by hand.
The verdict rests on simulated evidence and requires validation with real users before it is acted on.
The sequence, the cast, the machinery
Bokken executed the Design Thinking loop autonomously. Every actor below is journaled; persona contributions are simulated and labeled as such.
- 5 evidence items
- UI walkthrough
- 7 outcomes Ulwick-ranked
- 38 interpretations
- problem statement selected
- losers preserved with reasons
- 9 options, full lineage
- 3 firewalled lenses vote
- skeptic on record
- web research on the concept (sourced)
- 10 assumptions registered
- 2 artifacts hashed
- fresh firewalled panel
- 4/2/4 sup/con/untested
- verdict: iterate
- dossier (A/B/C)
- handoff specs
- this report (pptx + html)
Interview panel (6 personas)
Ideation panel (6 personas)
Test panel (6 personas — firewalled from the other panels)
System agents and models
| Agent | Did | Model calls |
|---|---|---|
facilitator | ran the stage machinery: programs, clustering, selection, register, fidelity, verdicts | |
ui-walker | walked the running app with a real browser; journaled observed facts and screenshots | |
concept-researcher | authorized deep web research on the selected concept, sources cited | |
convergence lenses | adversarial feasibility vs the codebase, independent RICE, outcome desirability | |
skeptic | mandatory on-record challenge before convergence closed | |
claude-fable-5 | routing classes served | 27 |
claude-haiku-4-5 | routing classes served | 9 |
claude-opus-5 | routing classes served | 11 |
What the run was grounded in
Brief
- Segment: daily pass commuters
- Segment: flexible pay-per-ride commuters
- Constraint: the nightly route optimization stays (it is the unit economics)
- Constraint: Spanish-first UI, plain language
- Constraint: no new hardware or driver-side changes
Tangible corpus
- code · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/repo
- metrics · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/kpis.csv
- discussion · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/interview_marta.md
- discussion · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/interview_diego.md
Evidence, on the record
Agents & activity
Javier (44, Zaragoza), Tomas (60, Zaragoza), Tomas (64, Badajoz), facilitator
10 research model call(s) on claude-fable-5. Kata moves: stage_contract.
Process
Corpus-calibrated interview program per segment; grounded persona interviews with citation-validated answers or honest abstention; functional UI walkthrough of the running app; JTBD desired-outcome derivation and per-persona Importance/Satisfaction scoring into a deterministic Ulwick opportunity ranking.
Output
20 evidence items, 1 abstentions, 7 ranked outcomes, UI walkthrough skipped.
1 question(s) were honestly abstained and carried as research debt.
Opportunity ranking (Ulwick: Opp = Importance + max(Importance − Satisfaction, 0))
Opportunity score per outcome (≥15 severely underserved, 12–15 underserved)
The problem statement, and why the losers lost
Agents & activity
facilitator
3 cognition model call(s) on claude-opus-5. Kata moves: hmw_reframe, stage_contract.
Process
Evidence clustered into insights tied to underserved outcomes; point-of-view candidates drafted and reframed (HMW) when solution-shaped; winner selected on evidence + opportunity coverage with losers preserved.
Output
One problem statement selected; losers preserved with reasons; opportunity coverage among the criteria.
How might we give riders an honest picture of tomorrow's commute before the booking deadline?
Flexible riders lack a fair price, pushing them to rail on marginal days
The Kata reframed solution-shaped statements 1 time(s) before selection.
9 options in, one concept out
Agents & activity
skeptic-agent, facilitator
4 challenge, 3 cognition, 9 extraction model call(s) on claude-fable-5, claude-haiku-4-5, claude-opus-5. Kata moves: stage_contract.
Process
Quota-driven work-alone divergence tied to outcome IDs with novelty monitoring; skeptic challenge on record; convergence through three firewalled lenses - adversarial feasibility vs the codebase (green/amber/red + first honest slice), independent RICE, outcome desirability.
Output
9 options with full lineage; one concept advanced; lens verdicts and dissent on record.
Dissent on record (feasibility)
Artifacts against the riskiest assumptions
Agents & activity
concept-researcher, ui-walker, facilitator
2 cognition, 2 generation, 2 research model call(s) on claude-fable-5, claude-opus-5. Kata moves: stage_contract.
Process
Assumption register built and risk-classified (impact x uncertainty); cheapest artifact set chosen against the riskiest assumption; artifacts generated and hash-journaled with assumption linkage.
Output
10 assumptions registered; 2 artifacts generated and hash-journaled.
Fidelity decision — why these artifacts
| Artifact | Kind | sha256 | Assumptions |
|---|---|---|---|
| artifacts/prototype/wireframe_html.html | wireframe_html | 7842679043e0 | 3 |
| artifacts/prototype/landing_copy.md | landing_copy | e3ffb3898c75 | 2 |
The web, on the record
Authorized deep research on the selected concept. Every signal carries its source; findings are reported evidence, not observation.
Competitors and prior art
| Who | What | Overlap |
|---|---|---|
| BusUp (demo data) | B2B commuter shuttles with contractual SLAs | partial - no consumer-facing window promise |
| Renfe Cercanias (demo data) | publishes real-time punctuality | the substitute riders defect to at 1.85 EUR |
Market signals
Differentiation risks
Open questions
The assumption register, scored
Agents & activity
Carmen (28, Zaragoza), Carmen (52, Zaragoza), Ramon (36, Madrid), facilitator
11 challenge model call(s) on claude-fable-5. Kata moves: loopback_proposal, stage_contract.
Process
Fresh firewalled panel evaluates the prototype against every register entry; quantified kill/iterate/proceed recommendation with loop-back proposal on contradiction.
Output
4 supported, 2 contradicted, 4 untested; recommendation: iterate.
Register outcome
Votes, challenge, dissent, iteration
Convergence lens votes
feasibility — 2 vote(s)
verdict green, effort S, first slice: push tomorrow's point+window at 21:00, read-only (no van switching yet) - The route engine already commits a plan nightly; surfacing it at 21:00 is exposure, not invention. The README confirms changes close at 21:00 - the data and the moment already align.
verdict red, effort L - A rider-voted pickup map fights the optimizer head-on; every vote becomes a constraint the engine must not break. Not honestly buildable as scoped.
viability — 2 vote(s)
RICE 4.2: reach 9 (every active rider gets the 21:00 push), impact 3 (attacks the top-ranked outcome, opp 16.0), confidence 0.7, effort 4.5 pw
RICE 0.6: monthly governance overhead for a fairness perception the receipts idea buys cheaper
desirability — 2 vote(s)
Directly serves the 16.0 outcome and Marta's exact words: 'I can plan around honesty'
Council governance is nobody's morning problem
The skeptic, verbatim
Dissent preserved on the decision
Facilitation (Kata) — every intervention journaled
| Move | Stage | Fired | Trigger / note |
|---|---|---|---|
stage_contract | empathize | yes | Opening empathize. Goal: understand the people in the problem space. Method: structured interviews with ladder |
stage_contract | define | yes | Opening define. Goal: frame the problem worth solving. Method: cluster evidence into insights, reframe, select |
hmw_reframe | define | yes | The current problem statement embeds a solution: "Build a night-before notification with tomorrow's window". R |
stage_contract | ideate | yes | Opening ideate. Goal: generate genuinely different options, then converge deliberately. Method: quota-driven d |
stage_contract | prototype | yes | Opening prototype. Goal: build the cheapest artifact that tests the riskiest assumption. Method: register assu |
stage_contract | test | yes | Opening test. Goal: score the assumption register against honest reactions. Method: fresh-panel evaluation (Do |
loopback_proposal | test | yes | A test result contradicts earlier work: the 'honesty reduces churn' assumption is contradicted by the same evi |
The Opportunity Solution Tree
Teresa Torres's discovery structure, read straight from the journal: outcome → opportunities (Ulwick-ranked) → the solution that advanced → its assumption tests.
Desired outcome (problem framed)
Daily pass commuters (46%→38% of riders) plan their mornings around a ±6-minute promise the product no longer keeps (71% compliance) nor admits to - the trust gap, not the delay itself, is driving the 345/month churnO3: Increase trust that the app's on-time indicator reflects the promise made to the rider, not the re-planned route - opportunity 16.0 (severely underserved)
O0: Increase the likelihood that the pickup window promised at 21:00 is the window that actually happens the next morning - opportunity 14.0 (underserved)
O1: Minimize the surprise of a pickup point that moved overnight - the rider knows the point and the walk before going to bed - opportunity 13.0 (underserved)
O2: Minimize the cost of a missed pickup - a same-morning recovery option instead of a 90-minute gap - opportunity 11.3 (moderate)
O5: Increase the perceived fairness of route optimization - riders understand why their pickup changed - opportunity 10.3 (moderate)
O6: Minimize the door-to-van walk variance across weeks - opportunity 9.7 (served)
Recommendation: iterate
Scored by a synthetic panel — requires validation with real users before acting.
Register: 4 supported · 2 contradicted · 4 untested of 10.
Loop-back proposal fired at test
What this run honestly did not do
Real-user questions the run refused to answer from its inputs — the honest to-do list for human research.
Every call journaled, every euro estimated
Estimated cost by model (USD)
| Model | Calls | Input tok | Output tok | Est. cost |
|---|---|---|---|---|
claude-fable-5 | 27 | 0 | 0 | $0.00 |
claude-haiku-4-5 | 9 | 0 | 0 | $0.00 |
claude-opus-5 | 11 | 0 | 0 | $0.00 |
| Total | $0.00 |
Companion records: dossier/dossier.md · dossier/dossier.json · journal.jsonl
What to do next, in order
Derived from journaled findings: broken features first, then the verdict's next step, then the top research debt.
- Address the test contradiction: A test result contradicts earlier work: the 'honesty reduces churn' assumption is contradicted by the same evidence that supports demand for honesty. Proposing we return to define rather than proceed on a weakened founda
- Recommendation is 'iterate': validate the supported assumptions with real users before building.
- Real-user research: Functional UI walkthrough of the running product
Specifications handed off
One sentence per spec; the full requirement text lives in the linked file.
handoff/openspec/changes/build-mvp-gallery/specs/honest-ontime-indicator/spec.md