Across four website design evals, model performance shifted as the prompts and eval setup changed. Claude Fable 5 and Claude Opus 5 both improved against Kimi K3, while GPT 5.6 Sol remained close to Kimi K3 across evals.

This made us curious: why did the Claude models perform better than Kimi K3 and GPT 5.6 Sol on certain prompts? What prompt strategies work best for each model family, if we want to optimize the design quality?
Specifically, the earlier evals asked models to build a landing page from a text brief, while the latest Human Creativity Benchmark (HCB) eval used a staged process with visual assets, mockup work, and refinement.
How the four website design evals were set up
The first three evals shared the same basic format. Each model received a written brief and produced one complete landing page, GPT 5.6 Sol ran through the API, and designers compared the outputs in blind general preference matchups. The prompt strategy, model lineup, and evaluator group changed between the three evals.
- Eval 1 used ten briefs ranging from loose prompts with an open visual direction to structured prompts with detailed requirements. The models were GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1.
- Eval 2 used ten briefs, including three loose prompts and seven structured prompts. The models were GPT 5.6 Sol, Claude Fable 5, Kimi K3, and Gemini 3.5 Flash.
- Eval 3 used ten briefs, with four loose prompts and six structured prompts. The models were GPT 5.6 Sol, Claude Fable 5, Kimi K3, and Claude Opus 5.
Eval 4, the Human Creativity Benchmark, changed the work more substantially. Instead of producing one complete page from a single brief, six models worked on three brands across three connected design stages. The models were GPT 5.6 Sol, Claude Fable 5, Claude Opus 5, Kimi K3, Gemini 3.6 Flash, and Muse Spark 1.1. Ideation established the page structure and content, mockup applied a visual system, and refinement made specific changes to the existing page. The eval also included more visual assets and ran GPT 5.6 Sol through the Codex CLI harness.
HCB asked designers about general preference, visual aesthetics, prompt adherence, and usability. This analysis uses only its general preference rounds, because that was the pairwise question shared with the first three evals.
GPT 5.6 Sol and Claude Fable 5 were the only models included in all four evals, giving us a consistent matchup for comparing the results. Kimi K3 and Claude Opus 5 provide additional context for whether the changes appeared across the wider model field.
GPT's lead disappeared in the HCB eval
GPT 5.6 Sol ranked first overall in each of the three earlier evals and beat Claude Fable 5 directly in all three. GPT won 53 of 90 matchups in Eval 1, 46 of 64 in Eval 2, and 70 of 100 in Eval 3, giving it head-to-head win rates of 59%, 72%, and 70%.
In the HCB eval, the earlier GPT-Fable result reversed. Claude Opus 5 and Claude Fable 5 ranked first and second overall, while GPT won 22 of 54 general preference matchups against Fable, dropping its head-to-head win rate from 70% in Eval 3 to 41% and putting Fable ahead at 59%.

Eval 1 was relatively close, while Eval 2 and Eval 3 showed much larger leads for GPT. HCB moved in the opposite direction, with GPT's win rate falling 29 points from Eval 3.
GPT stayed level with Kimi while the Claude models moved ahead
GPT 5.6 Sol's head-to-head result against Kimi K3 stayed almost unchanged across the three evals in which both appeared. GPT won 34 of 64 matchups in Eval 2, 55 of 100 in Eval 3, and 27 of 54 in the HCB eval, giving it win rates of 53%, 55%, and 50%.
Claude Fable 5 changed much more against the same opponent. It won 17 of 64 matchups against Kimi in Eval 2 and 33 of 100 in Eval 3, before winning 32 of 54 in HCB. Its win rate rose from 27% and 33% in the earlier evals to 59% in HCB. Claude Opus 5 followed the same pattern, winning 27 of 100 matchups against Kimi in Eval 3 and 34 of 54 in HCB, an increase from 27% to 63%.

GPT's steady result against Kimi makes a broad decline in its performance less likely. If GPT had simply performed worse in HCB, we might also expect its result against Kimi to fall rather than remain close to an even split. Fable and Opus, meanwhile, moved from trailing Kimi in the earlier evals to leading it in HCB. This suggests that the larger change was how well the Claude models performed under the HCB conditions, rather than GPT becoming weaker against every model.
GPT consistently performed well on loose briefs
We grouped the prompts in the first three evals by how much direction they provided. Loose briefs were shorter and more open, while structured briefs included more detailed requirements.
GPT beat Fable on loose briefs in all three evals. It won 29 of 36 matchups in Eval 1, 22 of 24 in Eval 2, and 27 of 40 in Eval 3, giving it win rates of 81%, 92%, and 68%. Although the size of the lead changed, GPT remained ahead on loose briefs every time.
The results on structured briefs were less consistent. GPT won 13 of 36 matchups in Eval 1, putting Fable ahead with a 64% win rate. GPT then won 24 of 40 structured matchups in Eval 2 and 43 of 60 in Eval 3, giving it win rates of 60% and 72%. Structured briefs moved from a Fable lead in Eval 1 to GPT leads in Eval 2 and Eval 3.

The HCB eval did not use the same loose and structured labels, but its prompts were more prescriptive than the earlier loose briefs. Even the ideation prompts specified the page hierarchy, required sections, CTAs, and reusable components. The mockup prompts defined the visual system, while the refinement prompts listed exact changes to make.
The move toward detailed, staged instructions may have contributed to GPT's lower win rate against Fable in HCB, because it moved the work away from the kind of brief where GPT had consistently led. Prompt structure was not the only part of the setup that changed, however. HCB also introduced a different workflow, more visual assets, new generated outputs, a different evaluator group, and a different way of running GPT, all of which could have affected performance as well.
GPT led during ideation and fell behind in the later design stages
Across the four models in this analysis, GPT 5.6 Sol led ideation with a 69% general preference win rate, Claude Opus 5 led mockup at 83%, and Claude Fable 5 led refinement at 57%. No model led every stage.

GPT's head-to-head result against Fable changed as HCB moved through its three connected stages. GPT won 11 of 18 ideation matchups, giving it a win rate of 61%. It then won 4 of 18 mockup matchups, a win rate of 22%, followed by 7 of 18 refinement matchups, a win rate of 39%, which put Fable ahead in both later stages.
These results show that the overall HCB rankings combined very different stage-level outcomes. However, they come from only three brands and one output from each model for every stage and brand, so they do not establish a stage pattern that we can expect to repeat for every landing page.
GPT led on loose briefs, while Claude models gained ground in staged visual work
Across the first three evals, GPT's most repeatable advantage against Fable came on loose briefs. GPT led Fable every time, with win rates ranging from 68% to 92%. Within these evals, GPT performed especially well when the brief was shorter, more open, and left more of the initial direction to the model.
The advantage moved toward the Claude models when the work shifted to HCB's staged workflow. GPT led Fable during ideation, but Fable moved ahead during mockup and refinement. Fable and Opus also moved from trailing Kimi in the earlier evals to leading it in HCB. Together, these results show that the Claude models performed better when the work placed more emphasis on applying a visual system, working with assets, and refining an existing direction.
For the landing-page work in these evals, the stronger model depended on where the task sat in the design process. GPT showed its clearest advantage when starting from an open brief, while the Claude models were stronger when the work involved applying a visual direction and refining it across stages.
What HCB should test next
The next version of HCB should keep the same ideation, mockup, and refinement workflow while expanding it across a wider range of brands, products, and visual directions. The prompt library should also include comparable tasks written with different levels of detail, allowing us to test loose and structured briefs within the same eval setup instead of comparing them across separate evals.
Each model should produce several outputs for every prompt rather than a single generation. Designers should continue making pairwise general preference selections while also rating each output individually, showing both which output was preferred and whether either one reached an acceptable level of quality.
HCB should also include the API and Codex CLI versions of GPT as separate participants working from the same prompts, assets, and workflow. Fable, Opus, and Kimi should remain as shared comparison models. This would help separate the effect of GPT's execution method while testing whether the model rankings hold across more brands, prompts, and generated outputs.
Limitations
The evidence is directional rather than a final model ranking. HCB produced 54 general preference votes in the GPT-Fable matchup, but those votes came from nine prompt-stage comparisons across three brands, with each comparison reviewed by six evaluators. The votes show how consistently those evaluators preferred the outputs they saw, but they do not represent 54 independent landing page tasks. The result also varied by brand: GPT's win rate against Fable was 17% on Arbórea Audio, 67% on Luma Iced Tea, and 39% on EPOCH Desk, so the overall HCB result depends partly on the three brands selected.
Evaluator agreement was modest across the four evals. The overall Krippendorff's alpha was 0.21, indicating limited agreement between evaluators. Two evaluators judging the same output pair agreed about 60% of the time. The evaluator groups changed between evals, which makes it difficult to separate evaluator calibration from the other changes in the eval setup.
Each model also produced one output for every prompt. The repeated Switchboard and Cobble briefs showed that a new output can produce a very different result, so one especially strong or weak generation can have a large effect on the observed model ranking.
Several parts of the eval setup changed together, including the task, prompts, assets, workflow, evaluator group, generated outputs, and the way GPT was run. Pairwise voting adds another limit, because it shows which of two outputs was preferred, but not whether either output reached an acceptable level of quality.
How we ran this study → Methodology
