Dataset·July 8, 2026·2 min read

Where AI product videos fall apart: 544 expert annotations

An open dataset of 544 timestamped notes from professional video editors on 15 product videos made by Google Veo 3.1, Adobe Firefly Video, and Grok Imagine.

Contra Labs
Contra Labs
July 8, 2026 · 2 min read

We asked professional video editors a simple question: where exactly do AI-generated product videos fall apart? This dataset holds 544 timestamped notes on fifteen videos made by Google Veo 3.1, Adobe Firefly Video, and Grok Imagine.

We gave all three models the same five prompts, with a focus on product shots. Professional video editors from the Contra network watched every clip and marked in detail what they saw as the positives and issues in these videos.

The professional editors were asked to label on five important video quality dimensions: Brand & Text Consistency, Material & Texture Realism, Camera Shot Adherence & Quality, Multi-Shot Cuts & Continuity, and Product Consistency. They also rated how much each issue matters, from High down to Low. What we collected is a close-up view of where the current frontier video models still struggle on real professional work.

Each annotation includes a start and end timestamp, a free-text comment, a dimension label, and a severity rating. This structure makes the data directly usable in several training contexts:

  • Multimodal reasoning mid-training: pair video segments with expert critique to teach models to reason about temporal consistency, physical plausibility, and visual quality in motion.
  • RLHF and reward modeling: use severity ratings as preference signals to train reward models that score generated video quality the way a professional would.
  • Fine-tuning video evaluation models: train automated critics or judges that can flag specific artifact types (text hallucination, texture instability, bad cuts) at the frame level.
  • Benchmarking and evals: the cross-model, same-prompt structure makes it a natural held-out test set for comparing new video models against a human expert baseline.

The dataset was collected in June 2026. It is released under the CC BY 4.0 license on Hugging Face.

Continue reading3 studies
All research
  1. July 16, 2026Research
    Where four AI models break when they build a landing page8 designers annotated 40 pages from GPT 5.6 Sol, Claude Fable 5, Grok 4.5, and Muse Spark 1.1, marking 754 failure points across 8 tags. Every model broke differently.Read
  2. July 16, 2026Benchmark
    GPT Image 2 won 41.9% of logo tournaments, and still only half its logos were client-ready10 brand designers judged GPT Image 2, Nano Banana Pro, MAI Image 2.5, and Meta Muse across 12 logo briefs. The winner is clear, and still only half its logos cleared the client bar.Read
  3. July 15, 2026Dataset
    How professional editors work: 234 annotated Premiere stepsA preview dataset of 234 annotated steps across 4 computer-use trajectories, recorded as professional editors built vertical social reels in Adobe Premiere Pro.Read