Research·July 24, 2026·5 min read

Five expert designers iterated on Google Stitch. None would show the result to a client.

Five expert product designers ran five Stitch iterations each on the same dashboard brief, tagging every issue at V1, V2, and V5. The problem count never went down, the fidelity never moved, and none of them would show the result to a client.

  1. 01Issues tagged barely moved: 42 at V1, 43 at V2, 44 at V5
  2. 02By V5, 19 problems were fixed, 23 remained, 13 new appeared
  3. 03Flagged inaccuracies came back as omissions and hallucinations
  4. 040 of 5 designers would show the result to a client
Contra Labs
Contra Labs
Research

Google Stitch launched with a bang. Threads went viral about it replacing entire design teams. "Claude Code + NEW Stitch 2.0 just changed how I design apps." "Google Stitch is Insane. You really don't need Framer or Webflow. It's so over for designers." Entire UIs shipped "in under 10 minutes. THREE prompts. That's it." You don't need Figma anymore, the workflow threads all say the same thing, You don't need a designer.

So we gave Stitch to five expert product designers, one shared brief for a data-heavy monitoring dashboard, five iterations each, every failure tagged, every reaction recorded.

Stitch creates as many problems as it fixes

After each iteration, designers tagged issues across five categories: prompt adherence, user experience, information hierarchy, visual design, and infographics. Every designer ran five iterations and scored the output at three points: V1 after the first generation, V2 after the second, and V5 after the fifth.

The total number of issues tagged barely changed across iterations. 42 tags at V1, 43 at V2, 44 at V5. But the problems themselves kept changing. By V5, 19 problems reported in V1 were fixed, while 23 remained, and 13 new ones had shown up. Going from V1 to V2, designers fixed 17 issues and 18 new ones appeared. From V2 to V5, they fixed 10 and 13 more appeared. Each round cleared about as many problems as it created, so the total stayed about the same while the specific problems kept changing round to round.

Issues tagged at V1, V2, and V5: fixed, remaining, and new.
Share

What gets fixed, and what shows up Instead

So which problems does iteration fix, and which does it introduce? The tags tell a clear story. The fixes a designer could point to and name usually landed. Ask for a specific color, a moved element, or a corrected label on one screen, and Stitch generally delivered it. What almost never happened was a fix landing clean, with nothing new breaking in its place.

The more common outcome was a swap: Stitch introduced a new problem when a designer asked it to fix an existing one. The new problem is almost never the same type as the old one. It's the overcorrection of it. Two designers flagged inaccuracies, and the inaccuracies became omissions: told it was getting things wrong, Stitch dropped the content instead. One designer flagged misleading, hard-to-read charts, and got overcrowded ones. Another flagged off-brand styling, and the restyle came back with broken typography and misaligned elements.

How flagged issue types shifted across iterations.
Share

On the first generation, four of five designers tagged inaccuracies. By the third iteration they had dropped to two designers. Omissions grew from two designers to four, as Stitch dropped requirements it had gotten right a round earlier. Hallucinations, which never came up at V1, showed up in two sessions, where Stitch added things nobody asked for. The corrections went in, and Stitch started leaving out what it used to include and added things it was never asked for.

Five designers, two kinds of sessions: one where things got better, and the other worse

Same tool, same brief, but the five sessions split into two experiences. Some got better as they went, and some got worse. You can see the difference in the tags and hear it in how the designers talked through it.

One designer had a rough start.

"Oh, no. The whole thing changed. Oh my God, the whole thing changed... I cannot find an undo button."

Then they changed how they prompted. Instead of describing the whole screen, they selected the one element they wanted to change and named the exact change they wanted. Their final prompt was one with really specific instructions. They asked to change the background, label, and icon color with the hex codes for each specified. Specificity produced the exact result they wanted. They were the only designer whose tag count fell over the course of the study, from 11 at V1 to 7 at V3, and at the end said, "I would say I am having fun, like, just exploring how this is coming together."

A single-element, hex-code-specific edit, and the result it produced.
Share

Another designer never wrote a vague prompt. Their revisions were lists with ten items at a time, each one naming a screen, an element, and the change they wanted. The detail, however, didn't lower their issue count. They tagged five issues in every round, a different five each time, as old problems got fixed and new ones appeared in their place. What it bought them was control. By The end, the designer said: "got this screen right finally... that's awesome."

A ten-item revision list naming a screen, an element, and a change per line.
Share

Other sessions went the other way. One designer asked for the same thing over and over: the same navigation on every screen. They asked in the second round, in 509 characters. Stitch did not do it. They asked again in the fourth round, in 280 characters. Stitch did not do it. They asked a third time, worn down to a single sentence: "the navigation menu still looks different across all pages." Stitch still did not do it. Their prompts got shorter as their frustration grew, and their tag count climbed from 14 to 18.

Three attempts at the same cross-screen navigation request.
Share

None of the Stitch outputs were client ready

Before they started iterating, the designers sorted Stitch's first output into one of two buckets: a wireframe (two of five) or a mockup (three of five). After five iterations, the buckets were exactly the same. Two wireframes, three mockups. None of the outputs moved closer to something a designer could hand to a client.

When asked whether they would present the output to a client, all five said no, including the designers who felt their session had gone well.

All five designers said no to presenting the output to a client.
Share

Even the designers whose sessions clearly improved landed in the same place as the ones who struggled.

What the five sessions teach about prompting Stitch

Put the prompts next to the results and a few rules stand out.

The opening prompt matters less than you would expect. The first prompts ranged from 33 characters to 2,300. Every one of them landed on a wireframe or a mockup with issues in most categories. What made the difference was everything that came after.

Ask for one screen at a time. The best moments in the study were all specific, single-screen requests. As one designer put it, "It does pretty good when you're iterating over the screen itself."

Be specific. The prompts that worked named the exact element and the exact change. A designer who asked for a precise color on a single element got exactly that, while vague direction like "ensure the consistency of UI design" did not.

Rephrase or work around instead of repeating yourself. One designer asked for consistent navigation across screens and got turned down three times in a row. When Stitch misses a cross-screen rule once, asking again more firmly does not improve the odds.

The designer asked: "I would like for you to check each selected page structure and menu and fix any navigation elements that look different on every page (….) It should be a single version across all selected screens."
Some examples of what would've worked instead: "use the navigation on the home page to replace the ones on the contact us page and gallery page", or "Copy the top navigation bar from the evals page and use it to replace the navigation on the settings page and the logs page."

Fix before you expand. Adding screens to a design that already has broken flows just spreads the problem. Every new screen a designer asked for came back with the same flow issues as the ones before it.

Methodology

Five expert product designers, sourced from Contra, each completed an independent session using Google Stitch. Every designer worked from the same brief, a developer-facing LLM observability dashboard. After each generation, designers tagged issues across five categories:

  • Prompt adherence
  • User experience
  • Information hierarchy
  • Visual design
  • Infographics

Each designer wrote their own opening prompt, then ran five iterations while narrating their thinking out loud. Sessions were screen-recorded end to end through Rollout, capturing every prompt exactly as typed. Output was evaluated at three points: V1 after the first generation, V2 after the second, and V5 after the fifth. At each point, designers classified the fidelity stage (wireframe, mockup, prototype, or production-ready), tagged issues across the five categories with the option to mark "no issues, just refinement changes," and, after the final generation, answered whether they would present the result to a client. Sentiment was scored from the think-aloud narration for each phase, and each selected tag, including written-in "other" tags, counts as one issue.

Limitations

Designers evaluated their own sessions, so the tags reflect first-hand judgment but carry no shared calibration step, which means one designer's tagged issue may be another's refinement. Tag counts are not comparable between designers, since some grade more strictly than others, so the claims here rest on trends within each designer rather than raw totals across the group.

With five designers and a single brief, each result reflects a small number of sessions, so a designer's outcome on any one session reflects their specific prompts and reactions, not a general rate. We treat these findings as directional rather than definitive, and place more weight on patterns that held across all five sessions, flat tag counts at the group level, no movement in fidelity stage, and a unanimous no on the client question, than on any single number.

The results also reflect one version of Stitch at one moment in July 2026. A newer version could behave differently, though the core pattern, that correcting one problem tends to surface another, is common to current generative tools and unlikely to disappear overnight.

How we ran this study → Methodology
Continue reading3 studies
All research
  1. July 23, 2026Benchmark
    Kimi K3 is a real rival to GPT 5.6 Sol on landing pagesKimi K3 arrived ranked first on Arena's Frontend Code Arena. Across 480 blind comparisons by 8 designers, it finished level with GPT 5.6 Sol on landing pages, and edged ahead on the detailed briefs.Read
  2. July 21, 2026Dataset
    What professional ad design looks like: 35 on-brief creativesAn open dataset of 35 finished social ad creatives, one for each of 35 synthetic brands across 4 industries, designed in Figma by professionals from the Contra network working from simulated client briefs.Read
  3. July 16, 2026Research
    Where four AI models break when they build a landing page8 designers annotated 40 pages from GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1, marking 754 failure points across 8 tags. Every model broke differently.Read