Google Stitch launched with a bang. Threads went viral about it replacing entire design teams. "Claude Code + NEW Stitch 2.0 just changed how I design apps." "Google Stitch is Insane. You really don't need Framer or Webflow. It's so over for designers." Entire UIs shipped "in under 10 minutes. THREE prompts. That's it." You don't need Figma anymore, the workflow threads all say the same thing, You don't need a designer.
So we gave Stitch to five expert product designers. Each worked from the same brief, a data-heavy monitoring dashboard, and ran five iterations. Every issue was tagged, every reaction recorded.
Every designer got something real out of their first prompt, in minutes. Each classified their own output: a wireframe (two of five) or a mockup (three of five). That part of the hype holds up. We wanted to see how far a designer could push it from there, and that's what five iterations did.
What Stitch does well
Stitch is at its best on ideation and small, specific edits. Two sessions show what that looks like.
One designer's first generation followed the prompt closely and came back with a design system with headline, body, label and color scheme for dark mode that just worked. Later iterations fixed a contrast issue they did not notice but appreciated, and filled a multi-series chart with plausible mock data. The designer's verdict was that "it actually figured out the chart better than I could have."
Another designer had a rough start, but quickly realized Stitch can handle specific, detailed requests well. Initially, they described desired changes to the whole screen, but got only undesirable results:
"Oh, no. The whole thing changed. Oh my God, the whole thing changed... I cannot find an undo button."
Then they changed how they prompted. Instead of focusing on the whole screen, they selected the one element they wanted to change and named the exact change they wanted. They asked to change the background, label, and icon color with the hex codes for each specified. Specificity produced the exact result they wanted. As a result, they were the only designer whose tag count fell over the course of the study, from 11 at V1 to 7 at V5, and at the end said, "I would say I am having fun, like, just exploring how this is coming together."

Both sessions worked the same way: generate, then make small, specific, single-screen edits.
Where Stitch struggles
The trouble started when designers asked for refinement: consistency across screens, layout fixes, and polish.
One designer asked for the same thing over and over: the same navigation on every screen. Their prompts shrank from 509 characters in the second round to 280 in the fourth, and finally to a single sentence: "the navigation menu still looks different across all pages." The issue persisted across iterations, and this designer's tag count climbed from 14 to 18.

Another designer took the opposite approach. Their revisions were ten-item lists, each naming a screen, an element, and the change they wanted. Each round Stitch fixed what they asked for, but produced new issues along the way. By the end: "I'd still call this a mock-up. I would present it as a very early draft prototype, but as a final version, I would not show this to anyone."

How far five iterations got
After each iteration, designers tagged issues across five categories:
- prompt adherence
- user experience
- information hierarchy
- visual design
- infographics
Designer scored the output at three points, after the first (V1), second (V2), and fifth (V5) generation.
Stitch was able to fix most specific designer requests, such as specific color changes, moving elements, or correcting labels. By V5, 19 of the problems tagged at V1 were fixed.
However, Stitch frequently created many new problems while fixing the old ones: from V1 to V2 designers fixed 17 issues and 18 new ones appeared, and from V2 to V5 they fixed 10 while 13 more appeared. The total issues stayed flat, while the specific problems kept changing round to round.

The new problem was usually an overcorrection of the old one. Two designers flagged inaccuracies, and the inaccuracies became omissions: told it was getting things wrong, Stitch dropped the content instead. One designer flagged misleading, hard-to-read charts, and got overcrowded ones. Another flagged off-brand styling, and the restyle came back with broken typography and misaligned elements.

This trade is why the fidelity classification never moved. The designers sorted Stitch's first output into a wireframe (two of five) or a mockup (three of five), and after five iterations the buckets were exactly the same. The first prompt delivered as much fidelity as five rounds of iteration and that's the strength. Stitch takes you from a blank page to a mockup in minutes, and the designer takes it from there.
What the five sessions teach about prompting Stitch
Put the prompts next to the results and a few rules stand out.
Use Stitch for ideation and mockups, not refinement. Stitch does really well at wireframing and mockups. Take the mockup into your own tools and push toward polish.
The opening prompt matters less than you would expect. The first prompts ranged from 33 characters to 2,300. Every one of them landed on a wireframe or a mockup with issues in most categories. What made the difference was everything that came after.
Ask for one screen at a time. The best moments in the study were all specific, single-screen requests. As one designer put it, "It does pretty good when you're iterating over the screen itself."
Be specific. The prompts that worked named the exact element and the exact change. A designer who asked for a precise color on a single element got exactly that, while vague direction like "ensure the consistency of UI design" did not.
Rephrase or work around instead of repeating yourself. One designer asked for consistent navigation across screens and got turned down three times in a row. When Stitch misses a cross-screen rule once, asking again more firmly does not improve the odds.
The designer asked: "I would like for you to check each selected page structure and menu and fix any navigation elements that look different on every page (….) It should be a single version across all selected screens."
Some prompts we'd expect to work, based on the patterns above: "use the navigation on the home page to replace the ones on the contact us page and gallery page", or "Copy the top navigation bar from the evals page and use it to replace the navigation on the settings page and the logs page."
Fix before you expand. Adding screens to a design that already has broken flows just spreads the problem. Every new screen a designer asked for came back with the same flow issues as the ones before it.
Methodology
Five expert product designers, sourced from Contra, each completed an independent session using Google Stitch. Every designer worked from the same brief, a developer-facing LLM observability dashboard. After each generation, designers tagged issues across five categories:
- Prompt adherence
- User experience
- Information hierarchy
- Visual design
- Infographics
Each designer wrote their own opening prompt, then ran five iterations while narrating their thinking out loud. Sessions were screen-recorded end to end through Rollout, capturing every prompt exactly as typed. Output was evaluated at three points: V1 after the first generation, V2 after the second, and V5 after the fifth. At each point, designers classified the fidelity stage (wireframe, mockup, prototype, or production-ready), tagged issues across the five categories with the option to mark "no issues, just refinement changes." Each selected tag, including written-in "other" tags, counts as one issue.
Limitations
Designers evaluated their own sessions, so the tags reflect first-hand judgment but carry no shared calibration step, which means one designer's tagged issue may be another's refinement. Tag counts are not comparable between designers, since some grade more strictly than others, so the claims here rest on trends within each designer rather than raw totals across the group.
With five designers and a single brief, each result reflects a small number of sessions, so a designer's outcome on any one session reflects their specific prompts and reactions, not a general rate. We treat these findings as directional rather than definitive, and place more weight on patterns that held across all five sessions.
How we ran this study → Methodology
