Best Performing Models10 models
Phase
LeaderboardPairwise comparison
Contra Labs
1602
Claude Fable 5.1
Claude Fable 5
Claude Opus 5
GPT 5.6 Sol (Codex)
Kimi K3
Qwen 3.8 Max
Gemini 3.6 Flash
Muse Spark 1.3
Deepseek V4 Flash
Muse Spark 1.1

General preference · Prompt adherence · Usability · Visual aesthetics

GPT-6 AstraNew entryLanding Pages09/09/2026

GPT-6 Astra: 2nd Overall, the Only Model That Finishes the Flow

  • GPT-6 Astra lands 2nd overall on our Human Creativity Benchmark leaderboard at medium thinking effort. It trails only Claude Fable 5.1, and stays ahead of GPT-5.6 Sol, which lands 4th.
  • Mockup is its strongest stage. Working from a brief, brand guidelines, and an approved Ideation screenshot, designers called its output "a refined premium product" and said its hero gradient "gives the impression that this is a luxury item." One rater added "the materials are beautifully presented and the color palette is nicely used."
  • Ideation is where it falls behind. Astra loses to both GPT-5.6 Sol and Claude Fable 5.1 at this stage. Working from the brief and brand guidelines alone, it follows the structure but misses the mood. Designers said "the layout includes all the sections listed in the prompt," but the hero still misses the tone the brief describes.
  • Complaints concentrate on the hero section and text size. One designer said "the overall text is very small, and some details are even smaller, which is an accessibility issue." Hero problems are worst at Refinement, where the prompt asks for a short list of targeted fixes and Astra changes things it wasn't asked to touch. In one brief it dropped the hero gradient entirely, and another designer said "the text on the hero lacks contrast and is hardly readable."
  • GPT-6 Astra is the only model in the study that built complete, working user flows. It shipped native dialog elements, a functional pre-order button that opens a cart with a running subtotal stored in local storage, a live count in the navigation bar, and a working remove button. Testers also noted it is bolder with color palettes and better at handling 3D scenes than its predecessor.
  • OpenAI says Astra "brings stronger visual judgment to front-end design," and one of its researchers claimed "significant progress on subjective domains like design aesthetics." Our testing found little improvement over its predecessor GPT-5.6 Sol on that front. Designers did praise its spacing, calling it "well applied" and saying the "page flows nicely, feels full but not crowded."
Mockup Elo ranking by model
GPT-6 Astra
1662
GPT-5.6 Sol
1552
Claude Fable 5.1
1489
Claude Opus 5
1489
Muse Spark 1.1
1308

Bradley-Terry Elo rating across all evaluation dimensions in Mockup. 208 retained comparisons. Baseline is 1,500.

Share
  1. 1Claude Fable 5.11602
  2. 2Claude Fable 51544
  3. 3Claude Opus 51543
  4. 4GPT 5.6 Sol (Codex)1532
  5. 5Kimi K31518
  6. 6Qwen 3.8 Max1511
  7. 7Gemini 3.6 Flash1474
  8. 8Muse Spark 1.31445
  9. 9Deepseek V4 Flash1438
  10. 10Muse Spark 1.11393
Methods & standardsSee how we score these models → Methodology

The world's leading independent human data & creative evaluation lab.

Powered by Contra.

The world's leading independent human data & creative evaluation lab.

Powered by Contra.