Video Arena — rank the best AI video models by human preference
The Video Arena is the live, blind, human-preference leaderboard for AI video generation models — the "renderer" category of world models. Two anonymous models receive the same starting image and action, each generates the next-world clip, and you vote for the more convincing result; the models are revealed only after you vote. Votes feed a regularized Bradley-Terry ranking (Elo-scaled, with confidence intervals) of the best image-to-video and text-to-video models.
How it works
- Pick a starting image and an action — or vote on an existing matchup.
- Two anonymous models generate the next-world clip from the same input.
- Vote blind for the better result, or call it a tie.
- Identities are revealed after the vote; the result updates the leaderboard.
What it ranks
Frontier image-to-video and text-to-video models — from labs such as Google (Veo), Kuaishou (Kling), OpenAI (Sora), and ByteDance (Seedance) — scored purely by what people prefer, not by automated metrics that correlate poorly with perceived quality. It is the first arena of WMArena, the World Model Arena.
See the live World Model Leaderboard, the ranking methodology, the World Model Arena, or what a world model arena is.