WorldSculpt decomposes video scenes into editable 3D meshes
Developed by Alaya Lab and the University of Tokyo, WorldSculpt is an open-source framework that reconstructs video scenes into hundreds of individually editable 3D object meshes. By conditioning single-object generative priors on multi-view inputs, it hallucinates watertight geometries for occluded objects to convert static scans into interactive environments.
Monolithic 3D representations like Gaussian splats look visually stunning, but they remain fundamentally non-interactive; modular, object-level meshification is the true missing link required to transform passive video captures into interactive gaming environments, AR/VR worlds, and robotics training simulators. By conditioning single-object canonical priors on multi-view camera poses, WorldSculpt bypasses the compute bottlenecks of full-scene 3D diffusion and resolves severe occlusion artifacts. This approach turns static 3DGS reconstructions into modular, physics-ready assets, though overall fidelity remains tied to the precision of upstream 2D and 3D perception pipelines.
DISCOVERED
1h ago
2026-09-13
PUBLISHED
1h ago
2026-09-13
RELEVANCE
AUTHOR
AI Search