Engineering · Inside Kibbo
Kibbo: from AI to a real-world adventure
Generating a story is only the beginning. The challenge is turning it into an experience that works beyond the screen.

Kibbo is a mobile app that turns exploring a city into a personalised adventure. The user provides a city, the kind of experience they want, the time available and how they plan to get around. Those choices become a route with checkpoints, a narrative and challenges tied to real places.
AI makes the stories and activities flexible: quizzes, observation riddles and photo challenges are among the checkpoint types the system supports. But the value of the product depends on something tangible: turning an interesting suggestion into an experience someone can use on their phone while out exploring.
The early technical challenges centre on that transition. I have organised the system around five pillars: grounding generation in the real world, defining clear contracts, involving the user in decisions, handling requests reliably and making results measurable.
In this article
01 / Real places
A stop needs to exist, not just sound plausible
A model can produce a convincing place name without identifying the right place. In Kibbo, that becomes a checkpoint: somewhere a person is expected to go. The pipeline therefore includes place discovery and verification through Google Places, followed by a shortlist supplied to generation.
When there are enough candidates, strict mode restricts checkpoints to the verified list and checks the returned place IDs. When there are not, the flow is less restrictive: grounding should not be presented as a universal guarantee. Coordinates are resolved, and missing coordinates are treated as an error rather than something to hide.
The challenge is balancing narrative freedom with geographical constraints. Finding a place on a map does not, on its own, establish that a route is walkable or accessible at all times.
02 / Explicit contracts
Model output is input that needs validation
The app cannot depend on free-form text that needs interpreting from scratch. Generation returns a structured adventure with defined fields and checkpoint types. The backend checks required fields, checkpoint order, duplicate place names and quiz structure.
Valid JSON is not the same as a valid adventure. A multiple-choice quiz, for example, needs four options and exactly one correct answer. Application checks enforce those rules, while the pipeline allows a limited number of retries with error feedback.
Responsibilities are separated too: the orchestrator distinguishes adventure creation, checkpoint validation, answer disputes and session resumption. The model contributes content; product state and rules remain the responsibility of the software.
03 / Human control
The user confirms before full generation
Before generating the complete adventure, the flow can ask for information and present a preview for confirmation. This is a human-in-the-loop step: the person takes part in the decision instead of simply receiving a final result.
That pause needs to exist in the backend. Sessions retain their context and state in the database; resuming checks the user, expiry and state version. A confirmation must refer to the right proposal, even when it arrives in a later request.
The challenge is not just displaying a Confirm button. It is modelling what happens before and after it, including cases where the session can no longer be used.
04 / Reliability
A double tap should not start two generations
Confirming a preview is a small but useful example. Two requests can arrive almost simultaneously. Disabling the button improves the interface, but the decisive check needs to happen on the server.
Kibbo uses an atomic conditional update on the run: only the request that successfully changes its status from waiting_human to running proceeds with generation. Other requests receive an already-processing response. This protects against concurrent confirmations; it is not a claim of exactly-once execution across the whole system.
This connects the mobile experience to persistent state management. React Native and Expo power the app; Supabase and Edge Functions provide the backend; PostgreSQL stores the data, and PostGIS supports geographical operations such as finding nearby adventures.
05 / Measurable quality
Observe runs and compare changes
When generation fails, it matters where: parsing, schema checks, coordinates, place constraints or external service calls. Run and tool-call logs, alongside token usage, estimated costs and latency, make those steps observable.
Generation also records a fingerprint containing versions and configuration. An eval suite calls the pipeline directly with fixed cases and compares different runs. Its checks cover structure, grounding, distance compatibility with the route, safety rules and gameplay.
Those scores are automated indicators, not proof that an adventure is enjoyable or safe in the real world. The challenge is using them to find regressions and compare changes without mistaking one good response for the quality of the whole product.
What makes the product real
The common thread is that AI does not remove the need for systems engineering: it makes it more visible. Generated content needs to pass through clear boundaries before it becomes something a user can rely on.
This first article covers Kibbo’s foundations. They are also starting points for a closer look at the project: how to select places, handle a confirmation and evaluate generation beyond a good first impression.