SpatialAxiom pushes open spatial reasoning forward
SpatialAxiom is an open-weight vision-language model family focused on 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding. Its 9B dense and 35B-A3B MoE models build on Qwen3.5 and support Transformers and vLLM.
SpatialAxiom makes spatial reasoning more practical for developers by pairing strong benchmark claims with familiar Qwen-based tooling. Its real test will be whether performance transfers from curated spatial evaluations to reliable perception in embodied systems.
- –Two model sizes cover local experimentation and higher-capacity serving
- –Full-parameter SFT keeps the architecture simple and provides a clean base for downstream fine-tuning or reinforcement learning
- –Training spans indoor scenes, egocentric views, and multi-camera settings
- –Reported results lead across several spatial benchmarks, but independent replication remains important
- –The CC BY-NC 4.0 license limits unrestricted commercial deployment
DISCOVERED
1h ago
2026-08-14
PUBLISHED
1h ago
2026-08-14
RELEVANCE
AUTHOR
gujiaqivadin