Apple-π benchmarks physical reasoning in video AI
Apple-π is a benchmark created by researchers at MMLab@NTU to test whether video generation and understanding models genuinely comprehend classical mechanics rather than relying on surface visual plausibility. Built around a three-stage scientific reasoning framework and the 400-video Orchard dataset, it diagnoses failure points across ten core classical mechanics tasks.
While current video generation models produce aesthetically impressive clips, they frequently rely on visual heuristics rather than real physical principles; Apple-π provides a much-needed diagnostic framework for true law-grounded physical intelligence.
• Evaluates models across three distinct reasoning stages—perception, formulation, and deduction—to pinpoint exact failure modes.
• Includes the Orchard dataset featuring 400 physics video scenarios across ten fundamental classical mechanics tasks.
• Establishes a benchmark requirement for world models to demonstrate grounded scientific reasoning instead of shortcutting visual output.
• Provides critical benchmarking infrastructure needed for trustworthy embodied AI and physical world simulation.
DISCOVERED
6h ago
2026-07-23
PUBLISHED
7h ago
2026-07-23
RELEVANCE
AUTHOR
_akhaliq