Back to home
Model Benchmarks
Real outputs instead of synthetic eval scores. Each benchmark gives every model the same prompts through the same pipeline, then shows you everything that came out, including the failures.
AI Video Generation
24 LLMs each generate 7 Remotion motion-design videos from identical prompts via FrameCall. The lineup includes GPT-5.6, Claude, Gemini 3.1 Pro, DeepSeek V4 and Kimi K2.6.
Watch the resultsOpenRouter Model Spend
Weekly USD spend per model on OpenRouter over the last 52 weeks, priced with time-accurate rates plus cache-adjusted estimates. Refreshes weekly.
See who pays the mostClaude Code Performance
Daily SWE-Bench-Pro pass rates for Claude Code on Opus 4.6, measured against a fixed baseline so 1-day, 7-day and 30-day regressions show up as they happen. Data from Marginlab, refreshed every 6 hours.
Check today's pass rateMore benchmarks coming. Have an idea for a model face-off? Tell us.