GenIA grounds SAM3D in observed views.
GenIA is an inference-time framework from Meta Reality Labs and the University of Tübingen that aligns SAM3D’s generative 3D prior with geometric and photometric observations. It reconstructs input-aligned objects from monocular images, multi-view captures, and videos without retraining the foundation model.
GenIA’s key contribution is treating alignment as a test-time problem rather than demanding a new foundation model. That makes SAM3D more useful for real capture workflows, though dynamic reconstruction still depends on external geometry estimation.
- –Visibility-aware attention, cross-view fusion, differentiable rendering, and optional refinement improve consistency with observed shape, pose, color, and detail.
- –The framework extends a single-view model to multi-view and dynamic inputs while keeping SAM3D frozen.
- –On dynamic sequences, GenIA reports substantially lower runtime than Lift4D and HiMoR, though it still adds significant inference overhead.
- –External ActionMesh geometry remains a dependency for dynamic scenes, creating failure modes around deformation and world-space placement.
- –The approach is promising for AR/VR, robotics, digital twins, and asset creation where faithful input alignment matters more than unconstrained visual plausibility.
DISCOVERED
1h ago
2026-10-11
PUBLISHED
1h ago
2026-10-11
RELEVANCE
AUTHOR
AI Search