Best Performing Models10 models
Phase
LeaderboardPairwise comparison
Contra Labs
1602
Claude Fable 5.1
Claude Fable 5
Claude Opus 5
GPT 5.6 Sol (Codex)
Kimi K3
Qwen 3.8 Max
Gemini 3.6 Flash
Muse Spark 1.3
Deepseek V4 Flash
Muse Spark 1.1

General preference · Prompt adherence · Usability · Visual aesthetics

Muse Spark 1.3New entryLanding Pages09/04/2026

Meta Muse Spark 1.3: 8th by professional design judgement, incredible cost effectiveness for frontier agent capabilities

  • Muse Spark 1.3 lands 8th overall, ahead of DeepSeek v4 and right behind Gemini 3.6 Flash. In our tasks, it costs about 23.8x less than Claude Fable 5.1 and 19.2x less than GPT-5.6 Sol while slightly edging out Fable 5.1 in the Mockup phase, where the models work from a brief, brand guidelines, and an approved screenshot of the Ideation phase.
  • The Ideation phase, working from a brief and the brand guidelines alone, is where it falls down. It wins only 10% of comparisons against Claude Fable 5.1 and GPT-5.6 Sol. Kimi K3 is the one open source frontier model it stays close to at this stage.
  • Designers' complaints concentrate on layout and spacing, and on vertical rhythm in particular. In one brief, designers complained "the content is pushed against the top edge" and "text under image is all too close together vertically.”
  • Against Muse Spark 1.1, it wins on every stage and every dimension we measured on the Human Creativity Benchmark: General Preference, Usability, Prompt Adherence, and Visual Appeal. The gap is widest in Refinement and on the Luma Iced Tea brief. In Refinement, 1.1 loses nearly every comparison.
  • While Muse Spark 1.3 follows a structure similar to its predecessor, Muse Spark 1.1, it uses better typography. One designer said "the type choices are restrained and professional." The pages also have better responsive layouts. One designer complained that "responsive design seems to be broken" on an output from Muse Spark 1.1.
  • Meta's claim of Muse Spark 1.3 having frontier-level performance is yet to be verified, and we will follow this up with another study once the max effort level model is out. Compared to its predecessor, 1.3 shows better performance in typography, color, and responsiveness, and we're excited to see what the Max model brings!
Generation Cost
Muse Spark 1.1
0.0093$
Muse Spark 1.3
0.0432$
Kimi K3
0.5248$
GPT 5.6-Sol
0.829$
Cladue Fable 5.1
1.0288$
Share
  1. 1Claude Fable 5.11602
  2. 2Claude Fable 51544
  3. 3Claude Opus 51543
  4. 4GPT 5.6 Sol (Codex)1532
  5. 5Kimi K31518
  6. 6Qwen 3.8 Max1511
  7. 7Gemini 3.6 Flash1474
  8. 8Muse Spark 1.3New1445
  9. 9Deepseek V4 Flash1438
  10. 10Muse Spark 1.11393
Methods & standardsSee how we score these models → Methodology

The world's leading independent human data & creative evaluation lab.

Powered by Contra.

The world's leading independent human data & creative evaluation lab.

Powered by Contra.