GPT-6 Astra: 2nd Overall, the Only Model That Finishes the Flow
- GPT-6 Astra lands 2nd overall on our Human Creativity Benchmark leaderboard at medium thinking effort. It trails only Claude Fable 5.1, and stays ahead of GPT-5.6 Sol, which lands 4th.
- Mockup is its strongest stage. Working from a brief, brand guidelines, and an approved Ideation screenshot, designers called its output "a refined premium product" and said its hero gradient "gives the impression that this is a luxury item." One rater added "the materials are beautifully presented and the color palette is nicely used."
- Ideation is where it falls behind. Astra loses to both GPT-5.6 Sol and Claude Fable 5.1 at this stage. Working from the brief and brand guidelines alone, it follows the structure but misses the mood. Designers said "the layout includes all the sections listed in the prompt," but the hero still misses the tone the brief describes.
- Complaints concentrate on the hero section and text size. One designer said "the overall text is very small, and some details are even smaller, which is an accessibility issue." Hero problems are worst at Refinement, where the prompt asks for a short list of targeted fixes and Astra changes things it wasn't asked to touch. In one brief it dropped the hero gradient entirely, and another designer said "the text on the hero lacks contrast and is hardly readable."
- GPT-6 Astra is the only model in the study that built complete, working user flows. It shipped native dialog elements, a functional pre-order button that opens a cart with a running subtotal stored in local storage, a live count in the navigation bar, and a working remove button. Testers also noted it is bolder with color palettes and better at handling 3D scenes than its predecessor.
- OpenAI says Astra "brings stronger visual judgment to front-end design," and one of its researchers claimed "significant progress on subjective domains like design aesthetics." Our testing found little improvement over its predecessor GPT-5.6 Sol on that front. Designers did praise its spacing, calling it "well applied" and saying the "page flows nicely, feels full but not crowded."
Bradley-Terry Elo rating across all evaluation dimensions in Mockup. 208 retained comparisons. Baseline is 1,500.