Best Performing Models9 models
Phase
LeaderboardPairwise comparison
Contra Labs
1603
Claude Fable 5.1
Claude Opus 5
Claude Fable 5
GPT 5.6 Sol (Codex)
Kimi K3
Qwen 3.8 Max
Gemini 3.6 Flash
Deepseek V4 Flash
Muse Spark 1.1

General preference · Prompt adherence · Usability · Visual aesthetics

Claude Fable 5.1New entryLanding Pages09/03/2026

Claude Fable 5.1: fewer generic pages, better storytelling, and the top spot on our Landing Page benchmark

  • Claude Fable 5.1 leads our Landing Page leaderboard overall and on every dimension we measured: General Preference, Usability, Prompt Adherence, and Visual Appeal.
  • Its lead is widest at Ideation, where it swept every comparison against Muse Spark 1.1. Professional designers praised its outputs in this phase on asset choice, visual consistency, and layout balance against the brief. One called it 'the perfect alliance between the prompt and the output.'
  • Fable 5.1 makes its mistakes in Mockup, where it works from a style kit and a screenshot of the Ideation output. Spacing and alignment drew the most complaints, followed by typography and contrast. Designers flagged an FAQ section that didn't line up with the sections around it, heading text with no line height set, and text set over imagery without enough contrast.
  • Fable 5.1 is also really strong at refining existing pages, rebuilding layout balance and completing the fixes designers ask for. One designer called the output 'balanced, engaging and sharp.'
  • Compared to Fable 5, we also noticed Fable 5.1 producing fewer generic outputs and better first-prompt results. Hero sections were more diverse and interactive, and the storytelling was stronger.
  • At High Effort, Fable 5.1 has a median generation time of 3 minutes 20 seconds per page, 21.1% faster than Claude Opus 5, while beating it in 60.8% of comparisons.
Win rate in the Ideation Stage · Landing Pages
Claude Fable 5.1
81.8%
Kimi K3
25%
Claude Opus 5
25%
Qwen 3.8 Max
22.7%
Muse Spark 1.1
0%

Claude Fable 5.1 leads Ideation, sweeping Muse Spark 1.1 on every comparison.

Share
  1. 1Claude Fable 5.1New1603
  2. 2Claude Opus 51537
  3. 3Claude Fable 51536
  4. 4GPT 5.6 Sol (Codex)1517
  5. 5Kimi K31509
  6. 6Qwen 3.8 Max1505
  7. 7Gemini 3.6 Flash1467
  8. 8Deepseek V4 Flash1431
  9. 9Muse Spark 1.11395
Methods & standardsSee how we score these models → Methodology

The world's leading independent human data & creative evaluation lab.

Powered by Contra.

The world's leading independent human data & creative evaluation lab.

Powered by Contra.