M
MesmerTools
APIsBenchmarksBlog
Back to home

Model Benchmarks

Real outputs instead of synthetic eval scores. Each benchmark gives every model the same prompts through the same pipeline, then shows you everything that came out, including the failures.

AI Video Generation

24 LLMs each generate 7 Remotion motion-design videos from identical prompts via FrameCall. The lineup includes GPT-5.6, Claude, Gemini 3.1 Pro, DeepSeek V4 and Kimi K2.6.

Watch the results

OpenRouter Model Spend

Weekly USD spend per model on OpenRouter over the last 52 weeks, priced with time-accurate rates plus cache-adjusted estimates. Refreshes weekly.

See who pays the most

Claude Code Performance

Daily SWE-Bench-Pro pass rates for Claude Code on Opus 4.6, measured against a fixed baseline so 1-day, 7-day and 30-day regressions show up as they happen. Data from Marginlab, refreshed every 6 hours.

Check today's pass rate
More benchmarks coming. Have an idea for a model face-off? Tell us.
M
MesmerTools

Discover the best AI tools and utilities

Directory

  • Browse Tools
  • Submit a Tool

Free Tools

  • JSON Formatter
  • AI Logo Maker
  • Token Price Calculator
  • AI Music Maker
  • Terminal Paste Cleaner
  • Claude Code Viewer

Resources

  • Blog
  • AI Video Benchmark
  • Developer APIs
  • Is it Peak Hours for Claude?
  • Claude Code Performance

© 2026 MesmerTools. All rights reserved.

Privacy PolicyTerms of Service