Benchmark·August 12, 2026·9 min read

Ad videos: six models across ideation, mockup, and refinement

The Human Creativity Benchmark moves to video: six models, three client campaigns, three phases each, judged head to head by working creatives. Seedance 2.0 won the set, and no model held the product together once it started moving.

  1. 01Seedance 2.0 won at 68.3% and led all three rated dimensions
  2. 02Omni Flash leads ideation at 67.5%, finishes last at 24.4%
  3. 03Grok runs the opposite arc: 36.9% ideation to 67.2% refinement
  4. 04Kling, the pilot's steadiest model, is now below 48% every phase
Contra Labs
Contra Labs
Research

Where the Human Creativity Benchmark left video

Earlier this year, we published the Human Creativity Benchmark (HCB), a benchmark that measures generative models on professional creative work. The framework had three commitments: diverse creative domains and output types, three phases per task (ideation, mockup, and refinement) following realistic creative workflows, and prompts elicited directly from creative professionals in Contra's network. The pilot ran 93 prompts across 80 evaluation sessions and produced roughly 15,000 individual judgments.

Expanding and Scaling the Human Creativity Benchmark

We are now expanding Human Creativity Benchmark from a pilot into a standing benchmark, one domain at a time. Ad images went first. This is about Video, starting with Motion videos, with more domains to follow.

Similar to Ad Images, Ad Videos raises the bar on exactly the failure the pilot flagged: motion, timing, and temporal consistency have to hold across every frame.

What we ran

Six video models, three client campaigns, three phases each.

  • Veo 3.1
  • Seedance 2
  • Gemini Omni Flash
  • Grok Imagine
  • Kling v3.0 Pro
  • Luma Ray 3

The campaigns are IARA Eau de Parfum, Mata Forma, and Ondes. Each brief carries the business context, the product, and the client needs.

Evaluators are graphic designers, ad designers, and creative directors, recruited from Contra. Six evaluators each judged all nine prompts, seeing all six outputs per prompt and comparing them head to head. 3,240 pairwise judgments in total, with a written rationale for every general-preference pick. Every pair is judged four times, once per dimension:

  • General preference. Which output has the stronger general quality? The exact question changes by phase: which would you develop further at ideation, which would you present to the client at mockup, which would you ship in the deliverable at refinement. Evaluators also add a written rationale for each preference selection.
  • Usability. Which output will you more likely use in a real client project? This measures the practical value of the AI generated output.
  • Prompt adherence. Which output more faithfully executes what the brief asked for?
  • Visual aesthetics. Which output looks better visually? Consider visual quality dimensions such as composition, hierarchy, color, etc.

Seedance 2.0 Ranks #1 by Professional Creatives

Across the full set, Seedance 2.0 was the clear winner, with a 68.3% win rate. Veo 3.1 is the runner-up, with Grok Imagine close behind.

Overall win rates across the full set.
Share

Seedance also ranks first on all three rated dimensions: usability at 69.6 percent, prompt adherence at 68.1 percent, and visual aesthetics at 70.4 percent. It beat every other model in their head-to-head records.

Breakdown by Creative Workflow Phases

Across the three phases, Seedance 2.0 is the only model that performs consistently, never dropping below 65.6 percent and finishing within eight points of the leader at every stage. Gemini Omni Flash starts first and drops significantly by refinement, while Grok Imagine and Veo 3.1 start in the bottom half and converge toward the top as the job narrows.

Win rates by phase.
Share

In the ideation phase, Seedance 2.0 and Gemini Omni Flash pull clear of the field at 71.1% and 67.5%, with the next-best model roughly 25 points back. The written rationales reward them for the same thing: outputs that make good starting points. One evaluator, picking Omni Flash, praised exactly the qualities a professional wants at the top of a job:

Option B hands down: Less contrast, higher dynamic range. More cinematic realism. Neutral grade gives breathing room for fine-tuning colors in post. Better skin-tones for the model's hands and body. More breathing room between shots. They are distributed more seamlessly across the 8 seconds. The IARA bottle has good Hero composition and the god rays effect works well. Cuts are gentler. Each action has enough time for the audience to appreciate what's going on.
Ideation phase win rates.
Share

In the Mockup phase, Veo 3.1 takes over with a 69.7 percent win rate, followed by Seedance 2.0 at 68.3 and Grok Imagine at 57.2 percent win rate. Veo's Prompt adherence mentions in this phase have the best ratio of any model, at 23 Praises to 5 Criticism. Gemini Omni Flash slips to 44.2 percent, and Kling V3.0 Pro and Ray V3.2 trail at 34.2 and 26.4 percent.

Mockup phase win rates.
Share

In the Refinement phase, Grok Imagine leads at 67.2 percent, just ahead of Seedance 2.0 at 65.6 percent and Veo 3.1 at 63.9 percent. Gemini Omni Flash, which finished within four points of the ideation winner, drops to last at 24.4 percent, 43 points below its ideation win rate.

Refinement phase win rates.
Share

Grok Imagine repeats a pattern similar to our pilot: last at ideation, strong at mockup, first at refinement. Gemini Omni Flash runs the opposite arc: it generates the best starting points in the field, then falls behind once the task becomes editing them in mockup and refinement.

Kling V3.0 Pro, the pilot's most consistent video model, falls behind, finishing below 48 percent in every phase. Seedance 2.0 takes its place as the most consistent model, staying within eight points of the leader in all three phases.

A Closer Look At Each Model

Seedance 2.0

Seedance 2.0 strengths and weaknesses.
Share

It draws the most praise and the least criticism in the set: 404 mentions of praise against 105 of criticism, nearly four mentions of praise for every one of criticism. And the praise is about craft: motion and camera work draws 68 mentions of praise against 7 of criticism, and lighting and color grading 63 against 4. One evaluator, comparing surf shots on the Ondes campaign, wrote:

A is much stronger. The wide angle close to surface shot behind the ankles is cinematic and immersive. Hero shot material! The sun rising up in the background gives a photorealistic beauty to the image. Water particles and beads are realistic… Very usable, high quality output.
Seedance 2.0 theme breakdown.
Share

In a study where realism drew more criticism than praise for five of six models, Seedance has 91 mentions of realism praise against 44 of criticism, the only model with a net positive, and the rationales say why:

I choose the output on the left because of it's pacing feels more intentional and aesthetic versus the output on the right… To my eye there is virtually no 'tells' in the first shot for the output on the left.

Its main weakness shows up when things move fast. Product fidelity is its narrowest margin, at 16 mentions of praise against 10 of criticism. On the Ondes campaign's carve-turn shot, an evaluator caught the product transforming mid-trick:

Seedance 2.0
As the rider does a 180 carve turn, the shape of the board changes. This is a classic hallucination as the model is trying to preserve the tail of the board the same at all times, it switches nose for tail which results in a break of realism and continuity.

Its losses follow the same pattern: exaggerated water physics, a surfer who "looks like he is floating over water, not gliding through it."

Veo 3.1

Veo 3.1 strengths and weaknesses.
Share

Product and reference fidelity leads its praise and it is the only model where that theme stands clearly above the field. When it wins, it wins on the hero product:

Definitely B (right)! Product's shape matches the reference. Push in camera movement is consistent. The composition of the hero product makes sense.
Veo 3.1 across the three phases.
Share

It has the widest phase-to-phase swing of any model that finishes near the top. At ideation it wins just 39.4 percent of its matchups, and aesthetic praise in its rationales barely outweighs criticism, 41 mentions against 38. By refinement that ratio reaches two to one, at 52 against 26, and its win rate reaches 69.7 percent in mockup and 63.9 percent in refinement, with the study's best prompt adherence numbers along the way. It is the only model whose aesthetic reception improves at every phase.

The output on the right is superior because of it's adherence to the prompt while maintaining a high degree of polish aesthetically… the output on the right creates a reason in the scene for the hue shift and tastefully executes it.
Craft and fidelity compared across models.
Share

Craft and fidelity turn out to be different skills in this study: Seedance draws sixteen lighting praises per criticism but only 1.6 on product fidelity, and Grok and Kling both draw more fidelity criticism than praise. Veo is the only model strong at both, at 4.6 praises per criticism on product fidelity and 3.5 on lighting and color. When it loses, it is usually on conservative camera work:

Option A (left) only does a dolly-in that's awkward and loses focus of the bag, also ending in a strange composition losing attention from the hero object.
Veo 3.1 prompt adherence record.
Share

It can also break scene logic: on the Mata Forma campaign, an evaluator picked its opponent because a bird flew "into the concrete room" through what should have been glass. Still, its prompt adherence record is the best in the set, at 61 mentions of praise against 15 of criticism.

Grok Imagine

Grok Imagine strengths and weaknesses.
Share

Grok divides evaluators on realism more than any other model, drawing 68 mentions of realism praise against 78 of criticism. On lighting and color there is no argument: 56 mentions of praise against 15 of criticism, its strongest theme. When it wins, the rationales explain why:

The right output has better dynamic range, photorealism and cinematic quality. The light leaks and god rays feel more natural and react better to the environment.

It does its best work late in the creative process, rising from 36.9 percent at ideation to 67.2 percent at refinement, the largest gain of any model, and the second study in a row where it has finished first in the final phase. It is also the model evaluators most often described as actually executing the correction:

The output on the right was selected because it more closely resembles the reference product… the camera movement and leaf animation follow the corrective prompt closer with the output on the right.

Its signature failure is the frame and composition. Composition and scale draw its most distinctive criticism: a surfboard "massive in relationship to the rider," and on the Ondes brief in the mockup where the product itself came apart:

The output on the right suffers from product persistence issues — the board appears to break into two pieces and the becomes a second board. Its confusing and disorienting to watch.

Gemini Omni Flash

Gemini Omni Flash strengths and weaknesses.
Share

It generates the best first drafts in the study and the worst last ones. Its ideation wins came from qualities professionals value at the start of a job: a neutral, gradeable image, gentle cuts, and unusual respect for technical specs.

Option B hands down: Less contrast, higher dynamic range. More cinematic realism. Neutral grade gives breathing room for fine-tuning colors in post… Cuts are gentler. Each action has enough time for the audience to appreciate what's going on.
Gemini Omni Flash theme breakdown.
Share

Its signature failure is coherence. It draws the most realism and artifact criticism of any model, 103 mentions against 64 praise, the widest gap in the set, and the highest temporal-consistency criticism by a wide margin. The rationales are a catalogue of things appearing, multiplying, and morphing:

The output on the left suffers from coherence issues with the leaves disappearing and appearing throughout the clip, refracting in the bottle without a matching leaf in the environment, as well as aspect ratio changes that are not prompted.

Its praise profile is otherwise average. It wins on nothing in particular and loses on glitches, and that is how a model goes from 67.5 percent at ideation to 24.4 percent at refinement.

Kling V3.0 Pro

Kling V3.0 Pro strengths and weaknesses.
Share

Kling's ledger tips negative overall: 194 mentions of praise against 228 of criticism. Its realism ledger nearly balances with 44 praises against 49 criticisms. Motion and camera work draws 44 criticisms against 30 praises, and pacing and editing runs 32 against 11, roughly three criticisms for every praise: shake, rushed cuts, choppy pacing. One evaluator chose it over Ray while conceding the handheld wobble, on the grounds that at least it was the kind of mistake a human crew could make:

The video shows an overall good execution of the prompt, despite being very AI-obvious and lacking that naturalness like the other video. however, the other video fails the biggest thing: adding the logo, which is against the prompt

When its physics land, they land. Despite not winning, Kling stays most competitive on the Ondes campaign with 61.1% win rate, after Seedance at 66.7%. One designer picking Kling over Seedance said:

The output on the left was selected despite not adhering to the prompt very well because the product looks similar in size and the surfer standing on it appears to be using it realistically. The output on the right suffers from strange water physics and the product does not appear to be accurate to the reference photos.

But the pilot's most dependable video model is now a bottom-half one, below 48 percent in every phase.

Ray V3.2

Ray V3.2 strengths and weaknesses.
Share

It finished last overall and last or second to last in every phase, and the rationales say why in unusually direct terms. Usability is its most distinctive criticism, by the largest margin of any theme for any model. Evaluators did not just prefer the other output, they said Ray's could not ship:

B on the other hand looks sloppy. The proportion of the rider to the board doesn't make any sense. He looks tiny! B shot is completely unusable.

Its strength is a distinctive look. An evaluator voting for it described a trade:

I feel it has much more personality than the one in the left, however, I feel the quality it's a little bit compromised… however, it has a more interesting concept and proposal in lighting and shadows.

Its wins were mostly because a professional could see how to salvage and use it: "as an editor, a blur mask can get rid of that somehow."

What each model is the best for

If you'd like to use one process across your creative progress, Seedance 2.0 is the strongest model per this comparison, with Veo 3.1 as the runner-up.

If you want the best of each phase, start ideation with Gemini Omni Flash or Seedance 2.0, build the mockup with Veo 3.1, and run refinement with Grok Imagine or Veo 3.1. Then route around problems as they come up.

If the product keeps drifting from the reference, go to Veo 3.1. If the clip needs a stronger color and mood, render the locked direction in Grok. If the physics of the scene are breaking, Kling is worth a look.

JobPickWhy
The overall pick, if you only want to use oneSeedance 2.0First on every rated dimension and never below 65.6 percent in any phase. The only model whose realism praise outweighs its criticism. Budget a check on the product under fast motion, since its one recurring failure is the product changing shape mid-move.
Generating direction ideas from a briefGemini Omni Flash, Seedance 2.0They finish 3.6 points apart at ideation, 25 or more clear of the field.
Building the chosen direction into a finished videoVeo 3.1Mockup is where Veo takes over, at 69.7 percent, with the study's best prompt adherence ratio (23 praises to 5 criticisms) and the strongest hero-product fidelity in the set.
Revising against a specific noteGrok ImagineRefinement leader at 67.2 percent, and the model evaluators most often described as actually executing the corrective prompt.
Mood and look explorationRay V3.2Its praise concentrates on lighting and color grading with annotators praising its personality.
How we ran this study → Methodology
Continue reading3 studies
All research
  1. August 7, 2026Research
    What stands between Opus 5 and client-ready pages: layout and readabilityNine designers annotated 20 landing pages one at a time, 738 notes in all. On Opus 5's pages the notes clustered in layout and readability, while brand fit and originality drew the fewest flags in the set.Read
  2. August 6, 2026Benchmark
    Six image models made real ads. Each broke differently.We scaled the Human Creativity Benchmark into a standing benchmark, starting with ad images: six models, three client campaigns, three phases each, judged head to head by professional creatives. Meta Muse Image won the set, and every model showed a signature failure.Read
  3. July 30, 2026Benchmark
    Claude Opus 5 nails the words and the mood. The finish is what holds it back.Across 600 blind comparisons and 400 write ups, designers liked the copy more than the rest of the page. Contrast, length and a few broken sections were what held it back.Read