4DAnyone turns casual video into 4D
4DAnyone reconstructs an animated human from a single uncalibrated monocular video by generating consistent multiview footage and converting it into a 4D Gaussian Splatting model. The research targets accessible volumetric capture without camera rigs, calibration, or studio hardware.
4DAnyone makes single-camera 4D capture look increasingly practical, but its biggest breakthrough is consistency at scale—not merely hallucinating another viewpoint.
- –Reference Context Packing keeps multiview conditioning bounded as generated views accumulate
- –Target Context Routing reduces structural drift between separately processed viewpoint groups
- –Skeleton conditioning provides a more reliable geometric signal than noisy dense depth or camera estimates
- –The pipeline could lower barriers for avatars, virtual production, games, telepresence, and robotics data
- –As with any monocular reconstruction system, hidden surfaces and fine details remain learned guesses rather than observed geometry
DISCOVERED
1h ago
2026-08-23
PUBLISHED
2h ago
2026-08-23
RELEVANCE
AUTHOR
AI Search
