Einsia Physics Benchmark Exposes Video AI Gaps
World Models' Last Exam in Physics tests eight video generators across 40 controlled tasks covering mechanics, optics, fluids, heat, electromagnetism, and surface tension. Across 1,280 videos, the top model scored just 57.76/100, showing that plausible visuals still frequently violate measurable physical relationships. [ArXiv](https://arxiv.org/abs/2610.08791)
This is a valuable corrective to video benchmarks that reward realism without testing whether anything obeys nature. The score is sobering, though it should be read as a diagnostic of observable physics—not a universal measure of world-model intelligence.
- –Uses task-specific measurements instead of relying only on human or vision-language-model judgments
- –Separates visual continuity from physical correctness, a crucial distinction for robotics and planning
- –Seedance 2.5 leads overall, but 57.76/100 leaves substantial room for improvement
- –Covers physics beyond mechanics, including reflection, melting, fluids, and electromagnetism
- –Observability failures can suppress physical scores, so passing the benchmark still does not prove full physical consistency
DISCOVERED
1h ago
2026-10-09
PUBLISHED
2h ago
2026-10-09
RELEVANCE
AUTHOR
lmoroney