Research·July 24, 2026·5 min read

Five designers, five rounds: Google Stitch proves its strength in ideation.

Five expert product designers ran five Stitch iterations each on the same dashboard brief. The first prompts delivered wireframes and mockups in minutes; five rounds of refinement never raised the fidelity.

  1. 01Every first prompt landed something real: 2 wireframes, 3 mockups
  2. 02Small, specific edits landed; one session's tags fell 11 to 7
  3. 03Refinement stalled: issues held at 42, 43, 44 across five rounds
  4. 04First prompt delivered as much fidelity as five rounds of iteration
Contra Labs
Contra Labs
Research

Google Stitch launched with a bang. Threads went viral about it replacing entire design teams. "Claude Code + NEW Stitch 2.0 just changed how I design apps." "Google Stitch is Insane. You really don't need Framer or Webflow. It's so over for designers." Entire UIs shipped "in under 10 minutes. THREE prompts. That's it." You don't need Figma anymore, the workflow threads all say the same thing, You don't need a designer.

So we gave Stitch to five expert product designers. Each worked from the same brief, a data-heavy monitoring dashboard, and ran five iterations. Every issue was tagged, every reaction recorded.

Every designer got something real out of their first prompt, in minutes. Each classified their own output: a wireframe (two of five) or a mockup (three of five). That part of the hype holds up. We wanted to see how far a designer could push it from there, and that's what five iterations did.

What Stitch does well

Stitch is at its best on ideation and small, specific edits. Two sessions show what that looks like.

One designer's first generation followed the prompt closely and came back with a design system with headline, body, label and color scheme for dark mode that just worked. Later iterations fixed a contrast issue they did not notice but appreciated, and filled a multi-series chart with plausible mock data. The designer's verdict was that "it actually figured out the chart better than I could have."

Another designer had a rough start, but quickly realized Stitch can handle specific, detailed requests well. Initially, they described desired changes to the whole screen, but got only undesirable results:

"Oh, no. The whole thing changed. Oh my God, the whole thing changed... I cannot find an undo button."

Then they changed how they prompted. Instead of focusing on the whole screen, they selected the one element they wanted to change and named the exact change they wanted. They asked to change the background, label, and icon color with the hex codes for each specified. Specificity produced the exact result they wanted. As a result, they were the only designer whose tag count fell over the course of the study, from 11 at V1 to 7 at V5, and at the end said, "I would say I am having fun, like, just exploring how this is coming together."

A single-element, hex-code-specific edit, and the result it produced.
Share

Both sessions worked the same way: generate, then make small, specific, single-screen edits.

Where Stitch struggles

The trouble started when designers asked for refinement: consistency across screens, layout fixes, and polish.

One designer asked for the same thing over and over: the same navigation on every screen. Their prompts shrank from 509 characters in the second round to 280 in the fourth, and finally to a single sentence: "the navigation menu still looks different across all pages." The issue persisted across iterations, and this designer's tag count climbed from 14 to 18.

Three attempts at the same cross-screen navigation request.
Share

Another designer took the opposite approach. Their revisions were ten-item lists, each naming a screen, an element, and the change they wanted. Each round Stitch fixed what they asked for, but produced new issues along the way. By the end: "I'd still call this a mock-up. I would present it as a very early draft prototype, but as a final version, I would not show this to anyone."

A ten-item revision list naming a screen, an element, and a change per line.
Share

How far five iterations got

After each iteration, designers tagged issues across five categories:

  1. prompt adherence
  2. user experience
  3. information hierarchy
  4. visual design
  5. infographics

Designer scored the output at three points, after the first (V1), second (V2), and fifth (V5) generation.

Stitch was able to fix most specific designer requests, such as specific color changes, moving elements, or correcting labels. By V5, 19 of the problems tagged at V1 were fixed.

However, Stitch frequently created many new problems while fixing the old ones: from V1 to V2 designers fixed 17 issues and 18 new ones appeared, and from V2 to V5 they fixed 10 while 13 more appeared. The total issues stayed flat, while the specific problems kept changing round to round.

Issues tagged at V1, V2, and V5: fixed, remaining, and new.
Share

The new problem was usually an overcorrection of the old one. Two designers flagged inaccuracies, and the inaccuracies became omissions: told it was getting things wrong, Stitch dropped the content instead. One designer flagged misleading, hard-to-read charts, and got overcrowded ones. Another flagged off-brand styling, and the restyle came back with broken typography and misaligned elements.

How flagged issue types shifted across iterations.
Share

This trade is why the fidelity classification never moved. The designers sorted Stitch's first output into a wireframe (two of five) or a mockup (three of five), and after five iterations the buckets were exactly the same. The first prompt delivered as much fidelity as five rounds of iteration and that's the strength. Stitch takes you from a blank page to a mockup in minutes, and the designer takes it from there.

What the five sessions teach about prompting Stitch

Put the prompts next to the results and a few rules stand out.

Use Stitch for ideation and mockups, not refinement. Stitch does really well at wireframing and mockups. Take the mockup into your own tools and push toward polish.

The opening prompt matters less than you would expect. The first prompts ranged from 33 characters to 2,300. Every one of them landed on a wireframe or a mockup with issues in most categories. What made the difference was everything that came after.

Ask for one screen at a time. The best moments in the study were all specific, single-screen requests. As one designer put it, "It does pretty good when you're iterating over the screen itself."

Be specific. The prompts that worked named the exact element and the exact change. A designer who asked for a precise color on a single element got exactly that, while vague direction like "ensure the consistency of UI design" did not.

Rephrase or work around instead of repeating yourself. One designer asked for consistent navigation across screens and got turned down three times in a row. When Stitch misses a cross-screen rule once, asking again more firmly does not improve the odds.

The designer asked: "I would like for you to check each selected page structure and menu and fix any navigation elements that look different on every page (….) It should be a single version across all selected screens."
Some prompts we'd expect to work, based on the patterns above: "use the navigation on the home page to replace the ones on the contact us page and gallery page", or "Copy the top navigation bar from the evals page and use it to replace the navigation on the settings page and the logs page."

Fix before you expand. Adding screens to a design that already has broken flows just spreads the problem. Every new screen a designer asked for came back with the same flow issues as the ones before it.

Methodology

Five expert product designers, sourced from Contra, each completed an independent session using Google Stitch. Every designer worked from the same brief, a developer-facing LLM observability dashboard. After each generation, designers tagged issues across five categories:

  • Prompt adherence
  • User experience
  • Information hierarchy
  • Visual design
  • Infographics

Each designer wrote their own opening prompt, then ran five iterations while narrating their thinking out loud. Sessions were screen-recorded end to end through Rollout, capturing every prompt exactly as typed. Output was evaluated at three points: V1 after the first generation, V2 after the second, and V5 after the fifth. At each point, designers classified the fidelity stage (wireframe, mockup, prototype, or production-ready), tagged issues across the five categories with the option to mark "no issues, just refinement changes." Each selected tag, including written-in "other" tags, counts as one issue.

Limitations

Designers evaluated their own sessions, so the tags reflect first-hand judgment but carry no shared calibration step, which means one designer's tagged issue may be another's refinement. Tag counts are not comparable between designers, since some grade more strictly than others, so the claims here rest on trends within each designer rather than raw totals across the group.

With five designers and a single brief, each result reflects a small number of sessions, so a designer's outcome on any one session reflects their specific prompts and reactions, not a general rate. We treat these findings as directional rather than definitive, and place more weight on patterns that held across all five sessions.

How we ran this study → Methodology
Continue reading3 studies
All research
  1. August 14, 2026Benchmark
    Claude models lead landing page design, but no model wins every design stageSix models built landing pages for three products across ideation, mockup, and refinement. Opus and Fable finished first and second overall, but a different model led each stage of the work.Read
  2. August 12, 2026Benchmark
    Ad videos: six models across ideation, mockup, and refinementThe Human Creativity Benchmark moves to video: six models, three client campaigns, three phases each, judged head to head by working creatives. Seedance 2.0 won the set, and no model held the product together once it started moving.Read
  3. August 7, 2026Research
    What stands between Opus 5 and client-ready pages: layout and readabilityNine designers annotated 20 landing pages one at a time, 738 notes in all. On Opus 5's pages the notes clustered in layout and readability, while brand fit and originality drew the fewest flags in the set.Read