Skip to content
cliphou.se
Cliphouse ResearchOctober 2026 / v1.1

Motion Graphics
Benchmark.

Find the right AI model for motion graphics
without testing every model yourself.

We give ten leading models the same three Remotion briefs and render every result. Compare the videos side by side, see which models work on the first attempt, and check the actual API cost before you choose.

01 / Comparison

One brief. Every model.

10 models · 2 independent runs

Storytelling, composition, and interface animation

12 seconds · 1080p · 30 fps · Silent
Product launch - full brief
Create a silent, 12-second product-launch video for "Relay", a fictional team-planning app. Use Remotion, 1920 x 1080, 30 fps.
From 0-3 seconds, introduce "Relay" and "Make room for focused work".
From 3-9 seconds, show "Plan together", "Protect focus time", and "See progress" through an animated product interface.
From 9-12 seconds, show "Start your next week with Relay". Keep the final message fully readable for at least two seconds.
Use the supplied fonts and palette. Choose the layout, visual style, and transitions. Include no additional claims, external assets, or audio. Return a complete composition that follows the supplied component contract.

Product launch, run 1. 10 completed entries.

AnthropicPremium

Claude Opus 5.5

API cost
$0.374
Generation
211.0s
Run details
Model ID
anthropic/claude-opus-5.5
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791473470-56iTa35pdSfCu38guL3O
Status
Passed on first attempt
Provider
amazon-bedrock
Generated
2026-10-08T15:31:10.508Z
Render time
47.6s
AnthropicMid-range

Claude Sonnet 5.5

API cost
$0.135
Generation
80.2s
Run details
Model ID
anthropic/claude-sonnet-5.5
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791474384-b1YYEhNEJVtkCP49xr4o
Status
Passed on first attempt
Provider
google-vertex/global
Generated
2026-10-08T15:46:24.224Z
Render time
212.1s
AnthropicBudget

Claude Haiku 5.5

API cost
$0.011
Generation
90.3s
Run details
Model ID
anthropic/claude-haiku-5.5
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791473327-4RQArKnAgnoFNuwUwbYh
Status
Passed on first attempt
Provider
google-vertex/global
Generated
2026-10-08T15:28:47.009Z
Render time
34.7s
OpenAIMid-range

GPT-6.1 Sol

API cost
$0.116
Generation
195.8s
Run details
Model ID
openai/gpt-6.1-sol
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791475260-TVsJveMp49wbEOd6AnOc
Status
Passed on first attempt
Provider
openai
Generated
2026-10-08T16:01:00.939Z
Render time
33.1s
OpenAIBudget

GPT-6 Luna

API cost
$0.0066
Generation
129.7s
Run details
Model ID
openai/gpt-6-luna
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791475475-yp2El122stCMuqOk5ZKr
Status
Passed on first attempt
Provider
openai
Generated
2026-10-08T16:04:35.984Z
Render time
34.4s
GoogleBudget

Gemini 3.8 Flash

API cost
$0.087
Generation
163.8s
Run details
Model ID
google/gemini-3.8-flash
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791475707-lGBxIH9syDuIvJU8sTgx
Status
Passed on first attempt
Provider
google-ai-studio
Generated
2026-10-08T16:08:27.843Z
Render time
33.8s
DeepSeekBudget

DeepSeek V4.1 Flash

Recovered / R1 · 1 additional attempt

API cost
$0.0068
Generation
51.9s
Run details
Model ID
deepseek/deepseek-v4.1-flash
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791484067-mR90dWIhvqONB97nu6X6
Status
Recovered in automated repair pass R1
Original outcome
Failed after 1 attempt. Invalid submission: Completion limit reached.
Original API cost
$0.0057
Additional repair cost
$0.0010
Total attempts
2
Provider
decart/fp4
Generated
2026-10-08T18:27:47.516Z
Render time
32.7s
Z.aiBudget

GLM 5.3 FlashX

Recovered / R1 · 1 additional attempt

API cost
$0.030
Generation
22.4s
Run details
Model ID
z-ai/glm-5.3-flashx
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791484229-6jyLrryIu1BtuiCKMLqq
Status
Recovered in automated repair pass R1
Original outcome
Failed after 1 attempt. Invalid submission: Expected ',' or '}' after property value in JSON at position 23223 (line 2 column 23222)
Original API cost
$0.021
Additional repair cost
$0.0091
Total attempts
2
Provider
z-ai/fp8
Generated
2026-10-08T18:30:29.372Z
Render time
32.7s
QwenOpen weights

Qwen3.8 2.4T A95B

API cost
$0.030
Generation
101.3s
Run details
Model ID
qwen/qwen3.8-2.4t-a95b
Reasoning effort
medium
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791476573-paPXDZ0Ahc1eIdUsTXWS
Status
Passed on first attempt
Provider
novita
Generated
2026-10-08T16:22:53.346Z
Render time
32.2s
Moonshot AIOpen weights

Kimi K3

API cost
$0.356
Generation
518.9s
Run details
Model ID
moonshotai/kimi-k3
Reasoning effort
high
Timing source
OpenRouter generation_time, successful attempt only
Generation ID
gen-1791476771-hzcy498AYiscSSPjZhR7
Status
Passed on first attempt
Provider
decart/mxfp4
Generated
2026-10-08T16:26:11.447Z
Render time
36.4s

A curated selection across price tiers, not a popularity ranking. Empty entries are untested, not failed.

Generation is OpenRouter's reported time for the successful attempt that produced the video. It excludes previous failed attempts and local rendering. Missing OpenRouter timings are marked unavailable, never estimated. Timing records

Recovered entries show total API cost, including original attempts. The additional repair cost is itemized in run details. Recovery means the code compiled and rendered. It is not first-attempt success.

02 / Methodology

Same starting point.
No hand-picked winners.

Protocol v1.1

This tests language models writing motion-graphics code, not native text-to-video models. Every request goes through OpenRouter. The generated code is rendered with the same Remotion environment.

What stays the same

Three fixed briefs, two independent runs, one component contract, the same local fonts and palette, and a 32,000-token completion limit that includes reasoning. No browsing, external assets, or existing Cliphouse compositions.

What gets another attempt

Only a failed compile or render gets one repair request, with the error log. Successful videos receive no creative revision. Original attempts and repair costs are retained. Transport errors are recorded separately.

Automated recovery / R1

Only original failures receive up to three additional automated repair requests. Keep the same model, pinned provider, reasoning effort, brief, fonts, palette, and renderer. Return complete TSX directly; a single code fence or valid source JSON wrapper is accepted without editing the code. Supply the previous response and complete compiler/render diagnostics. Stop at the first successful render. No human code edits or creative feedback. Preserve original outcomes and report additional attempts and costs separately. Recovery uses a 64,000-token limit and direct TSX output instead of the original JSON-only format. These are additional repair results, not replacement first attempts.

Repair results · Repair protocol · Original results

What we disclose

The exact model ID, provider, request, reasoning setting, source code, OpenRouter generation time, and API cost. Generation time comes from the successful request's generation_time field in OpenRouter's generation metadata, converted from milliseconds to seconds. Providers are pinned and fallback is disabled. Up to four API requests overlap; renders remain sequential. Reasoning labels are not equivalent compute budgets across models.

What we mean by motion graphicsExisting Cliphouse template. Not a benchmark submission.
10models
3fixed briefs
2runs per brief

Both runs are published, including failures. Two runs are exploratory evidence, not a statistically definitive ranking.

Protocol revision: A Haiku setup pilot used 13,309 reasoning tokens and was truncated at the original 16,000-token limit. Protocol 1.1 uses 32,000 tokens for all 60 comparable runs. The pilot cost $0.008 and remains in the spending record, outside the comparison.

Repair-feedback limitation: In this pilot, compile-failure repairs received a stack trace without the detailed TypeScript diagnostics. Treat repair outcomes as provisional. First-attempt outcomes are unaffected.

03 / Editions

A record, not a moving target.

New model versions join the live comparison. Monthly reports preserve the tested versions, protocol, and results. Changes to the test receive a new protocol version.

Live comparisonLatest benchmark entries

Model selection checked October 8, 2026. Availability and prices can change.

Sources: OpenRouter model catalog, usage rankings, and provider routing.