Back to blog
Matt Clannachan 7 min read

How AI Actually Changed Our Mockup Workflow (Not the Way We Expected)

AI workflow transformation in mockup generation

When we started building Flowstep in early 2025, the working assumption was that AI would primarily speed things up. The hypothesis was: if a designer can generate a plausible screen in seconds instead of hours, the overall design cycle gets faster. That turned out to be true, but it was not the interesting result. The interesting result was something we did not anticipate at all.

What AI generation actually changed, more than anything else, was how quickly we could determine whether a brief was specific enough to act on.

The brief-quality problem we did not know we had

Before we built Flowstep, our design process used written briefs as inputs to manual mockup work. A product manager or a designer would write something like "settings page with notification preferences" and work from there. The ambiguity in that brief was not particularly visible because the person receiving it was a human with context. They would fill in the gaps from their knowledge of the product, make reasonable assumptions, and ask questions when something was genuinely unclear.

The first thing that became apparent when we started running those same briefs through AI generation was that "settings page with notification preferences" generates twenty different screens depending on what the system decides to infer. A human designer makes one coherent set of inferences. An AI-assisted system shows you the space of possible interpretations, not a single one.

That was jarring at first. We spent considerable time trying to make the generation more deterministic, tighter, less varied. But we were solving the wrong problem. The variation was not a bug in the generation. It was diagnostic information about the brief. A brief that produces ten wildly different interpretations is not a specific brief. A brief that produces a narrow range of closely related outputs is one the system and the designer agree on.

What changed: brief authoring became a skill we started training

Once we understood the diagnostic value, the workflow shifted. Instead of treating brief quality as a prerequisite, we started treating the generation step as a brief-quality check. You write a brief, you run it, you look at the output and ask: is this the range of screens I intended? If not, what did I leave unspecified?

The specific gaps that tend to surface are predictable once you have seen them a few times. User permissions and access levels: a brief that says "admin panel" but does not specify which role is viewing it generates outputs that differ substantially on what controls are visible. State handling: a brief that describes a form without specifying the validation state generates screens with no error states, which means the developer will ask about those states later, at a point where they are more expensive to address. Component hierarchy: a brief that says "primary action" but does not clarify whether it is a standalone button or part of a button group generates both, and neither is wrong from the brief's perspective.

None of these gaps are obvious before you see a screen generated from an underspecified brief. They become obvious immediately after. The generation step creates feedback that a written review of the brief typically does not.

The feedback loop we built around this

We formalized this into the current Flowstep workflow in a fairly simple way. When a generation runs, the system flags the elements in the output that were resolved by inference rather than by explicit brief specification. If a component's state was not named in the brief, the output marks it. If a layout decision was ambiguous, it marks that too.

The flags are not errors. They are decision points. The designer can review the flagged elements and either accept the inference (which then gets recorded so it applies consistently in subsequent generations from the same brief) or update the brief with an explicit specification.

Over several iterations, the brief acquires specificity it did not have at the start. The generation output narrows. The designer is not spending time in the generation tool repairing outputs, they are spending time in the brief itself, which is closer to the actual design decisions.

The part that surprised us most

The deepest change was not in how fast screens got made. It was in when disagreements surfaced. Before this workflow, misalignment between what a PM meant and what a designer built typically became visible at design review, after the work was done. With this workflow, the first generation attempt often produces something that makes the misalignment immediately visible, before significant work is invested.

Consider a brief for a checkout flow. The PM writes "billing details step with address form and order summary sidebar." A designer reads this, makes assumptions about which fields are required, whether the sidebar is collapsible, where validation messages go. A generation from the brief shows all of those decisions as outputs. If the PM looks at the generated screen and says "that's not what I meant by order summary sidebar," the disagreement surfaces in the first review of a generated screen rather than after a week of careful mockup work.

We did not build Flowstep as a communication tool. We built it as a generation tool. The communication improvement emerged from the brief-validation loop as a side effect. We now think it is one of the more useful properties of AI-assisted generation that is not talked about in the usual framing of "faster" and "more output".

What did not change

The design judgment required after generation did not become less demanding. A generated token-aligned screen is a better starting point than a blank canvas, but it is still a starting point. The visual hierarchy, the information architecture, the specific treatment of edge cases: all of that still requires a designer making considered decisions.

We want to be specific about this because the "AI generates designs automatically" framing sets incorrect expectations. The generation step produces a structurally correct, token-aligned initial frame from a well-specified brief. What it does not produce is a final design. The value is in the starting point quality and the brief-validation loop. The design work is still design work.

Where we are now

By April we had built enough of the brief-feedback loop into the core workflow that it was no longer an optional step. It runs on every generation and the outputs include explicit brief-quality feedback alongside the generated screens. We track which brief elements were resolved by inference and which by explicit specification, and that information is visible in the generation output.

The original hypothesis, faster screen generation, is also true. But the more durable change is that the quality bar for what counts as a "ready" brief has moved upstream. A brief that would have been considered complete before now gets annotated with its inferential gaps before any design work starts. That has been the more meaningful shift.

More from the blog

Turn your next brief into screens today.

Token-aligned screen generation, starting free. No card needed.

Start free