Lanzadera is an on-demand commuter shuttle for Spanish suburbs whose core covenant is a +/-6-minute pickup window.
What this run concluded
Every number on this page is derived from the append-only Journal — nothing was written by hand.
The verdict rests on simulated evidence and requires validation with real users before it is acted on.
The sequence, the cast, the machinery
Bokken executed the Design Thinking loop autonomously. Every actor below is journaled; persona contributions are simulated and labeled as such.
- 13 evidence items
- 4 features UI-tested
- 7 outcomes Ulwick-ranked
- 41 interpretations
- problem statement selected
- losers preserved with reasons
- 9 options, full lineage
- 3 firewalled lenses vote
- skeptic on record
- web research on the concept (sourced)
- 10 assumptions registered
- 2 artifacts hashed
- fresh firewalled panel
- 4/2/4 sup/con/untested
- verdict: iterate
- dossier (A/B/C)
- handoff specs
- this report (pptx + html)
Interview panel (6 personas)
Ideation panel (6 personas)
Test panel (6 personas — firewalled from the other panels)
System agents and models
| Agent | Did | Model calls |
|---|---|---|
facilitator | ran the stage machinery: programs, clustering, selection, register, fidelity, verdicts | |
ui-walker | walked the running app with a real browser; journaled observed facts and screenshots | |
concept-researcher | authorized deep web research on the selected concept, sources cited | |
convergence lenses | adversarial feasibility vs the codebase, independent RICE, outcome desirability | |
skeptic | mandatory on-record challenge before convergence closed | |
claude-fable-5 | routing classes served | 33 |
claude-haiku-4-5 | routing classes served | 9 |
claude-opus-5 | routing classes served | 12 |
claude-sonnet-5 | routing classes served | 9 |
What the run was grounded in
Brief
- Segment: daily pass commuters
- Segment: flexible pay-per-ride commuters
- Constraint: the nightly route optimization stays (it is the unit economics)
- Constraint: Spanish-first UI, plain language
- Constraint: no new hardware or driver-side changes
Tangible corpus
- code · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/repo
- running app · file:///Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/app/index.html
- metrics · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/kpis.csv
- discussion · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/interview_marta.md
- discussion · /Users/maglionejuanmartin/code/bokken/src/bokken/demo/fixtures/interview_diego.md
Evidence, on the record
Agents & activity
Javier (44, Zaragoza), Tomas (60, Zaragoza), Tomas (64, Badajoz), ui-tester, ui-walker, facilitator
1 cognition, 16 research, 9 sidekick model call(s) on claude-fable-5, claude-opus-5, claude-sonnet-5. Kata moves: stage_contract.
Process
Corpus-calibrated interview program per segment; grounded persona interviews with citation-validated answers or honest abstention; functional UI walkthrough of the running app; JTBD desired-outcome derivation and per-persona Importance/Satisfaction scoring into a deterministic Ulwick opportunity ranking.
Output
28 evidence items, 0 abstentions, 7 ranked outcomes, UI walkthrough done.
Opportunity ranking (Ulwick: Opp = Importance + max(Importance − Satisfaction, 0))
Opportunity score per outcome (≥15 severely underserved, 12–15 underserved)
The product, exercised first-hand
A browser walkthrough of the running app — screenshots and facts are observed evidence, not simulation.
The switch responds instantly and confirms inline ('mañana te recoge la van de las 07:20') without leaving the screen - exactly the one-tap recovery the wide-window warning needs.
Day and van selectors plus one confirm tap; the confirmation promises the 21:00 notice, which is the right moment to promise.
The card shows a green 'Ventana cumplida' tick directly above a notice that the arrival alert went out 9 minutes after the promised window - the indicator measures the dawn re-plan, not the 21:00 promise. This is the exact 'green tick that lies' from the Marta interview, reproduced in the product's own UI.
The opt-in checkbox exists, but its state does not visibly persist when switching views, so whether the preference sticks cannot be confirmed from the UI alone.
Full heuristic review (verbatim)
## What works at first contact
- Tomorrow's card leads with the two facts riders plan around - the window (07:38-07:46) and the walk (3 min) - and the one-tap van switch confirms inline without a page change.
- Booking is three taps end to end, and the confirmation promises the 21:00 notice: the product already knows its moment of truth.
## Findings
1. **Ayer card -> green 'Ventana cumplida' beside a 9-minutes-late notice -> this is the trust-breaking contradiction driving churn (Marta: 'a green tick that lies') -> measure the tick against the 21:00 promise, never the dawn re-plan.** The fine print admits the re-plan baseline; honesty is one comparison swap away.
2. Ajustes -> the 21:00 opt-in checkbox does not visibly persist across views -> riders who opt in may silently stay unnotified -> persist and echo the state ('te avisaremos a las 21:00').
3. Ajustes -> the 89-euro pass banner shows regardless of riding pattern -> 2-3 day riders read it as mispricing (Diego) -> gate the banner on trips/week.
Verdicts: 2 works · 1 broken · 1 unclear. The broken finding is the same insight the interviews surfaced - the UI reproduces the dishonest indicator faithfully.
The problem statement, and why the losers lost
Agents & activity
facilitator
3 cognition model call(s) on claude-opus-5. Kata moves: hmw_reframe, stage_contract.
Process
Evidence clustered into insights tied to underserved outcomes; point-of-view candidates drafted and reframed (HMW) when solution-shaped; winner selected on evidence + opportunity coverage with losers preserved.
Output
One problem statement selected; losers preserved with reasons; opportunity coverage among the criteria.
How might we give riders an honest picture of tomorrow's commute before the booking deadline?
Flexible riders lack a fair price, pushing them to rail on marginal days
The Kata reframed solution-shaped statements 1 time(s) before selection.
9 options in, one concept out
Agents & activity
skeptic-agent, facilitator
4 challenge, 3 cognition, 9 extraction model call(s) on claude-fable-5, claude-haiku-4-5, claude-opus-5. Kata moves: stage_contract, timebox_pivot.
Process
Quota-driven work-alone divergence tied to outcome IDs with novelty monitoring; skeptic challenge on record; convergence through three firewalled lenses - adversarial feasibility vs the codebase (green/amber/red + first honest slice), independent RICE, outcome desirability.
Output
9 options with full lineage; one concept advanced; lens verdicts and dissent on record.
Dissent on record (feasibility)
Artifacts against the riskiest assumptions
Agents & activity
concept-researcher, ui-walker, facilitator
2 cognition, 2 generation, 2 research model call(s) on claude-fable-5, claude-opus-5. Kata moves: stage_contract.
Process
Assumption register built and risk-classified (impact x uncertainty); cheapest artifact set chosen against the riskiest assumption; artifacts generated and hash-journaled with assumption linkage.
Output
10 assumptions registered; 2 artifacts generated and hash-journaled.
Fidelity decision — why these artifacts
| Artifact | Kind | sha256 | Assumptions |
|---|---|---|---|
| artifacts/prototype/wireframe_html.html | wireframe_html | 7842679043e0 | 3 |
| artifacts/prototype/landing_copy.md | landing_copy | e3ffb3898c75 | 2 |
The web, on the record
Authorized deep research on the selected concept. Every signal carries its source; findings are reported evidence, not observation.
Competitors and prior art
| Who | What | Overlap |
|---|---|---|
| BusUp (demo data) | B2B commuter shuttles with contractual SLAs | partial - no consumer-facing window promise |
| Renfe Cercanias (demo data) | publishes real-time punctuality | the substitute riders defect to at 1.85 EUR |
Market signals
Differentiation risks
Open questions
The assumption register, scored
Agents & activity
Carmen (28, Zaragoza), Carmen (52, Zaragoza), Ramon (36, Madrid), facilitator
11 challenge model call(s) on claude-fable-5. Kata moves: loopback_proposal, stage_contract.
Process
Fresh firewalled panel evaluates the prototype against every register entry; quantified kill/iterate/proceed recommendation with loop-back proposal on contradiction.
Output
4 supported, 2 contradicted, 4 untested; recommendation: iterate.
Register outcome
Votes, challenge, dissent, iteration
Convergence lens votes
feasibility — 2 vote(s)
verdict green, effort S, first slice: push tomorrow's point+window at 21:00, read-only (no van switching yet) - The route engine already commits a plan nightly; surfacing it at 21:00 is exposure, not invention. The README confirms changes close at 21:00 - the data and the moment already align.
verdict red, effort L - A rider-voted pickup map fights the optimizer head-on; every vote becomes a constraint the engine must not break. Not honestly buildable as scoped.
viability — 2 vote(s)
RICE 4.2: reach 9 (every active rider gets the 21:00 push), impact 3 (attacks the top-ranked outcome, opp 16.0), confidence 0.7, effort 4.5 pw
RICE 0.6: monthly governance overhead for a fairness perception the receipts idea buys cheaper
desirability — 2 vote(s)
Directly serves the 16.0 outcome and Marta's exact words: 'I can plan around honesty'
Council governance is nobody's morning problem
The skeptic, verbatim
Dissent preserved on the decision
Facilitation (Kata) — every intervention journaled
| Move | Stage | Fired | Trigger / note |
|---|---|---|---|
stage_contract | empathize | yes | Opening empathize. Goal: understand the people in the problem space. Method: structured interviews with ladder |
stage_contract | define | yes | Opening define. Goal: frame the problem worth solving. Method: cluster evidence into insights, reframe, select |
hmw_reframe | define | yes | The current problem statement embeds a solution: "Build a night-before notification with tomorrow's window". R |
stage_contract | ideate | yes | Opening ideate. Goal: generate genuinely different options, then converge deliberately. Method: quota-driven d |
timebox_pivot | ideate | yes | Idea novelty has decayed (17% new against a floor of 20%). Proposing we move from divergence to convergence. |
stage_contract | prototype | yes | Opening prototype. Goal: build the cheapest artifact that tests the riskiest assumption. Method: register assu |
stage_contract | test | yes | Opening test. Goal: score the assumption register against honest reactions. Method: fresh-panel evaluation (Do |
loopback_proposal | test | yes | A test result contradicts earlier work: the 'honesty reduces churn' assumption is contradicted by the same evi |
The Opportunity Solution Tree
Teresa Torres's discovery structure, read straight from the journal: outcome → opportunities (Ulwick-ranked) → the solution that advanced → its assumption tests.
Desired outcome (problem framed)
Daily pass commuters (46%→38% of riders) plan their mornings around a ±6-minute promise the product no longer keeps (71% compliance) nor admits to - the trust gap, not the delay itself, is driving the 345/month churnO3: Increase trust that the app's on-time indicator reflects the promise made to the rider, not the re-planned route - opportunity 16.0 (severely underserved)
O0: Increase the likelihood that the pickup window promised at 21:00 is the window that actually happens the next morning - opportunity 14.0 (underserved)
O1: Minimize the surprise of a pickup point that moved overnight - the rider knows the point and the walk before going to bed - opportunity 13.0 (underserved)
O2: Minimize the cost of a missed pickup - a same-morning recovery option instead of a 90-minute gap - opportunity 11.3 (moderate)
O5: Increase the perceived fairness of route optimization - riders understand why their pickup changed - opportunity 10.3 (moderate)
O6: Minimize the door-to-van walk variance across weeks - opportunity 9.7 (served)
Recommendation: iterate
Scored by a synthetic panel — requires validation with real users before acting.
Register: 4 supported · 2 contradicted · 4 untested of 10.
Loop-back proposal fired at test
What this run honestly did not do
No open research debt was journaled in this run.
Every call journaled, every euro estimated
Demo session: the usage below is an illustrative live-run profile journaled by the scripted provider — nothing was charged and no network call was made. A first real run on your own product typically lands at $20-35 list price.
Estimated cost by model (USD)
| Model | Calls | Input tok | Output tok | Est. cost |
|---|---|---|---|---|
claude-fable-5 | 33 | 298,800 | 73,500 | $7.74 |
claude-haiku-4-5 | 9 | 32,400 | 8,100 | $0.07 |
claude-opus-5 | 12 | 134,400 | 56,400 | $2.16 |
claude-sonnet-5 | 9 | 25,200 | 15,300 | $0.39 |
| Total | $10.36 |
Companion records: dossier/dossier.md · dossier/dossier.json · journal.jsonl
What to do next, in order
Derived from journaled findings: broken features first, then the verdict's next step, then the top research debt.
- Fix (Indicador de puntualidad de ayer): The card shows a green 'Ventana cumplida' tick directly above a notice that the arrival alert went out 9 minutes after the promised window - the indicator measures the dawn re-plan, not the 21:00 promise. This is the exact 'green tick that lies' from the Marta interview, reproduced in the product's own UI.
- Address the test contradiction: A test result contradicts earlier work: the 'honesty reduces churn' assumption is contradicted by the same evidence that supports demand for honesty. Proposing we return to define rather than proceed on a weakened founda
- Recommendation is 'iterate': validate the supported assumptions with real users before building.
Specifications handed off
One sentence per spec; the full requirement text lives in the linked file.
handoff/openspec/changes/build-mvp-gallery/specs/honest-ontime-indicator/spec.md