The Intent Model
Write it down at the beginning, before anything can drift.
A BRD is a business requirements document, the 40-page thing a client hands you that’s supposed to pin down what you’re building. It describes the system in enormous detail and commits to none of it. Hand one to 5 people, or to 5 LLM sessions, and you get 5 different apps, and every one of them can be defended against the text.
That’s what a vehicle booking system for port logistics looked like at the start: about 40 pages, 4 actor types, 37 business rules, and 3 state machines (shipment, slot booking, pickup execution) that overlapped in everyone’s head. A state machine is just the list of conditions a thing can be in and the moves it’s allowed to make between them. So on any enterprise build I take on now, one file gets written before a line of code. A structured JSON contract that fixes what the app is supposed to do. I call it the intent model.
Compressing the BRD
Before anything gets designed or generated, the document gets squeezed into one typed artifact with 6 sections: actors, entities, journeys, business rules, constraints, and open questions. I chose that set and have been developing it since February. No single field in it is new. DDD gave me entities, use cases gave me actors and journeys, state machines and requirements engineering gave me the rest. Putting all of them in one file is the part I own, and everything downstream derives from that file instead of going back to re-read the BRD.
The squeezing is where the ambiguity dies, because the format has nowhere to hide a contradiction. Best example from that build: an HBL (House Bill of Lading, the document a shipment travels under) turned out to have 2 lifecycles running at once. Where the container physically is, on_vessel → at_wharf → in_yard → unpacked → collected. And who’s responsible for collecting it, unassigned → assigned → delegated → booked. Both true at the same time, for the same HBL. The BRD blended the two the whole way through and never mentioned it. Untangling them took 3 meetings with the client. Those meetings were the actual work. The file is just where the answer gets to live.
An entity, once it’s in:
{
"name": "HBL",
"key_fields": [
{ "name": "hbl_number", "type": "id", "required": true },
{ "name": "customs_status", "type": "enum", "enum_values": ["held", "cleared"] }
],
"lifecycle_states": ["on_vessel", "at_wharf", "in_yard", "unpacked", "collected"],
"state_machine": {
"transitions": [
{
"from": "unpacked",
"to": "collected",
"trigger": "driver_pickup",
"guard": "customs_status == cleared"
}
]
}
}
The guard is the load-bearing part. A guard is a condition that has to be true before a move is allowed. “No collection until customs clears” used to be a sentence somewhere in the middle of the BRD. Now it’s a named condition on one specific transition. Every screen with a pickup button disables it until customs clears, the generated tests assert exactly that, and no future session has to rediscover the rule by rereading 40 pages.
The Memento analogy
In Memento, Leonard can’t form new memories, so he runs his investigation on Polaroids, notes, and tattoos. The film is about how that system betrays him. He writes everything down after the fact, under stress, with none of the context attached, and by the end he’s editing his own notes to steer his future self. His last act in the film is writing down something he knows is false, because he’d rather have a mission than the truth.
Every LLM session has Leonard’s condition. It wakes up with no memory of yesterday, and it fills every gap you leave with the most statistically ordinary thing that fits. People call this hallucination. Under-specification is the better word, because the gaps get filled from training data instead of from your project.
The intent model is the tattoo, with 2 corrections Leonard never made. The writing happens at the beginning, while the client is still in the room to say “no, delegated and booked are different states,” instead of being reconstructed afterwards from a fading memory of a meeting. And the notes validate: a tattoo can say anything, but an intent file has to pass a schema check, and every change lands in a changelog with a version, an author, and a date. The analogy breaks in one place worth naming. Leonard could never lay out every Polaroid at once and check them against each other. The pipeline does exactly that, on every commit.
Why JSON
Half of it is about ambiguity and survival. Prose carries its ambiguity all the way to the model, JSON forces me to resolve it first, and models are less likely to “helpfully” rewrite a JSON file than a markdown one, so the contract tends to survive contact.
The other half is enums. An enum is a fixed list of allowed values, nothing else permitted. A prose spec can say a conflicting edit “should generally defer to the server” and move on. The JSON version has to pick one of last_write_wins | server_wins | prompt_user | merge, and there’s no enum value for “generally.”
For the business rules I borrowed EARS, a notation developed at Rolls-Royce while they were analysing the airworthiness rules for a jet engine control system. Every rule has to take one of 5 shapes:
| Pattern | Template |
|---|---|
| Event-driven | WHEN [event] THE SYSTEM SHALL [behavior] |
| State-driven | WHILE [state] THE SYSTEM SHALL [behavior] |
| Unwanted behavior | IF [condition] THEN THE SYSTEM SHALL [behavior] |
| Optional feature | WHERE [feature enabled] THE SYSTEM SHALL [behavior] |
| Ubiquitous | THE SYSTEM SHALL [behavior] |
It reads stiff on purpose. If a rule can’t be phrased as one of these, nobody has actually decided it yet, so it goes into open_questions[] with an ID and a reason and waits for a human answer instead of getting a statistical one. That section turned out to be the rarest thing in the file. I later checked the 6 sections against about 30 requirements frameworks, and only 5 of them carry any equivalent of a slot for the undecided.
Enterprise SLAs
“Enterprise SLA-grade” sounds like a sales word, and it isn’t. It’s a specific bar written into the contract. On one enterprise email build, a full mail, calendar, and contacts suite, the obligations include a 99% crash-free rate and a 2-hour response window on severity-1 defects. You don’t hit numbers like that by generating happy paths and patching whatever breaks in QA. Every failure behavior has to be written down before anyone can hit it.
So the email app’s intent model carries contracts the booking system never needed. Every mutation, and there are 264 of them, declares what it does optimistically, what it does when that fails, and how it settles a conflict, all before any code exists:
{
"id": "archive_thread",
"applies_to": "mail_thread",
"trigger": "swipe_right",
"optimistic_behavior": {
"action": "remove_from_active_list",
"feedback": "trigger_light_haptic_and_show_toast"
},
"rollback_behavior": {
"action": "restore_to_original_index",
"feedback": "show_error_toast_with_retry"
},
"conflict_resolution": "last_write_wins",
"offline_queueable": true
}
Every screen, all 97 of them, declares an empty state, a loading state, and an error state before it’s allowed to exist. The offline-capable ones declare a stale state on top, which is what you see when you’re reading cached mail with no signal. And every requirement from the client’s original list sits in a triage map with a status, including the ones we’re not building. Descoped items stay in the file marked descoped, so a scope audit can prove nothing was dropped quietly, and the generation step refuses to build from anything unsigned.
This is what “unbreakable no matter what you throw at it” actually cashes out to. Swipe to archive in a dead zone with no signal, and the thread comes back at its original index with a retry toast, because that exact behavior was agreed on months before the screen existed. Nobody is being vigilant in the moment. The failure modes were enumerated while everyone was still calm, and the pipeline won’t generate a screen that ignores them.
Two ways to use it
The booking system and the email app use the intent model in nearly opposite ways, which taught me more about the idea than either project did alone.
| Booking system | Email app | |
|---|---|---|
| Main job | Get client and team to agree before dev starts | Drive code generation through the whole build |
| Shape | One TypeScript file, 157 KB | Ten JSON files, one per domain, about 2 MB |
| Carries | 4 actors, 17 entities, 16 journeys, 37 rules, 50 open questions | All of that plus 97 screens and 264 mutation contracts |
| Reviewed by | Humans: per-section approve or dispute, client sign-off | Tooling: schema lint and audit on every commit |
| After shipping | Retires; code becomes the truth, a drift report keeps score | Stays the source of truth; generated files carry provenance headers |
On the booking system the model is a consensus instrument. It has its own small review app where stakeholders walk through each section and approve or dispute it, and the effect on the schedule was blunt: roughly 3 weeks getting humans to agree, then about a week and a half of shipping screens.
On the email app it’s a machine contract, split into bounded contexts (mail, auth, calendar, and 7 more) so a session working on compose loads intent-mail.json and nothing else, never the whole contract. Every generated file opens with a header like // @generated-from: intent-mail.json#archive_thread, so when something looks wrong, the argument happens at the model and the fix propagates, instead of getting patched in place and forgotten.
The 2 files diverge on how far each project lets the model reach rather than on how big it is, and that turns out to change what belongs in a spec at all. I’ve gone into that separately in what actually goes in a spec.
What the model feeds
Agreement is the model’s first job. Robustness comes from what gets generated off it: Zod schemas for every entity, typed API stubs, SQL migrations, screen scaffolds, TanStack Query hooks (one per declared mutation), a design handover doc filtered down to what designers need. XState skeletons for the state machines get generated once and are hand-owned afterwards, because regenerating a state machine a human has spent weeks hardening is how you lose the hardening. Even the BRD flipped direction on the booking system. The model generates the document now, not the other way around.
On the booking system the export crosses a repo boundary. The model compiles to a single intent-contract.ts that lands in the portal codebase with a do-not-edit header, and a drift script compares the app’s real types against it, field by field, and writes a report of what’s missing.
The visual half gets the same discipline through design tokens. One theme defines every color, radius, and spacing value. Components consume tokens. Figma receives the same tokens through a plugin and is allowed to recompose and refine, never to mint a new color or a new component. The flow only goes one way:
TweakCN theme editor
└─ base.css every color, radius, spacing token
├─ component registry ──→ generated screens
└─ Figma (same tokens) ──→ visual refinement only
(The email app pushes the token idea further. It’s white-label, so a tenant supplies one accent hex and the theme system expands it into a full ramp at runtime. One authored color, whole product themed.)
Put the 3 together and each one covers the others’ blind side. The intent model pins down what the app does, the token system fixes what it looks like, and the generation pipeline is the only road from either of them into the code, stamping provenance on everything it touches. A session boxed in on all 3 sides can still be wrong. It can’t be quietly wrong, and quiet wrongness is what actually kills enterprise builds.
What might make this obsolete
Everything above is scaffolding around 2 model weaknesses: sessions forget, and gaps get filled with confident averages. Both are shrinking with every release. Context windows keep growing, memory keeps improving, and I can picture the version where a model reads the whole BRD directly, holds all of it, asks its own 3 meetings’ worth of clarifying questions, and builds the app with the logic in its head. No JSON intermediary, no guardrails, no me. If that arrives, a good chunk of this article turns into advice on organizing floppy disks, and I’d rather say so here than pretend the method is eternal.
The part I don’t expect to age out was never for the model. The 50 open questions in the booking system’s file weren’t things an AI failed to understand. They were things no human had decided yet. The sign-off block exists because stakeholders disputed rules and someone had to be on record resolving them. When a client asks in month 4 whether drivers were ever supposed to see pricing, the answer comes back with a rule ID, a version, and a date. Models are getting better at remembering. Getting 5 humans to agree on what’s true, and being able to prove it later, was never the model’s job, and writing it down at the beginning is still the only fix I know.
Meanwhile the tooling keeps me honest about my own discipline. While writing this piece I checked the version numbers: the booking system’s intent model is at 0.14.0, the contract exported into the portal came from 0.11.0, and the drift report was last generated against 0.7.1. The system for catching drift had drifted. It took 3 numbers and about a minute to see that, which is most of the argument for writing them down in the first place.